Insights · Research note ·

Prompt less, check more: programmatic rules for reliable customer-service agents

An AI agent can follow a rule in one reply and break it in the next. We tested code checks on proposed messages and actions, with specific feedback when a check fails. Carefully designed checks reduced detected violations in our tests. Poorly designed checks made agents less effective.

Hivemind Labs · Research note

Detected violations in the experiments

Scroll sideways to read the chart.

Without gatesCorrected gates + specific feedback (post-hoc)BENCHMARK · conversations with a detected violationτ²-bench retail · flash90 conversations4.4%0.0%τ²-bench retail · Haiku30 conversations6.7%0.0%0%2%4%6%8%PRODUCTION REPLAY · 74 previously rejected turns · flashWithout gates56.8%3 gates + specific feedback51.4%Same outputs, echo removed25.7%0%20%40%60%Echo removal was applied before scoring; it was not implemented in the sending pipeline.
Figure 1. Detected violations in benchmark conversations and previously rejected production turns. Benchmark results use the corrected gates (post-hoc). Production results show both unmodified outputs and outputs with the echo removed before scoring. The panels use different percentage scales; definitions are in sections 2.4 and 3.

0 detected / 120

The corrected nine-gate setup had zero detected violations across 120 test conversations on two models. The gates were revised after the original results and tested on the same tasks.

Completion uncertain

There was no statistically clear change in task completion. The intervals still allow meaningful losses or gains. Measured flash agent cost per conversation was similar.

55% fewer after echo removal

In 74 previously rejected turns, violations fell from 56.8% to 51.4% with three gates, or 25.7% after removing the echo before scoring. That automatic fix was not implemented in the sending pipeline.

Feedback setup matters

The tested specific-feedback setup performed better than generic retry in several comparisons. Other prompt details also varied, so the explanation’s separate effect remains unresolved.

1. Introduction

A gate is a code check before an agent sends a message or takes an action. It checks the proposed output against a rule. If the check fails, the agent receives specific feedback about what went wrong and what would satisfy the rule, then tries again.

For example, an agent says an appointment is booked, but no booking tool has succeeded. A gate catches the unsupported claim. The feedback asks the agent to complete the booking or correct its reply before it reaches the customer.

Our production system has used a check layer since June 2026. It runs string checks first, tool-call checks next, and model-graded rules where code cannot express the requirement. Each rule has a configured consequence: send, hold for approval, or suppress. The same rules support tests before publication and checks on live turns.

This study tests deterministic checks only. It asks how they affect violations, task completion and cost, and how the tested feedback setup compares with a generic retry. In deployment, unresolved failures can be held for a person. In these experiments, the final attempt was released even if a check still failed, so the remaining violations could be measured.

2. Method

2.1 Benchmarks and corpora

τ²-bench retail [1] is a customer-service benchmark with deterministic tools over a simulated store, a 1,200-word written policy, and a simulated user driven by a language model. Reward is programmatic: the final database state, the required actions, and the information communicated to the user. We use a fixed 30-task subset of the base split sampled with seed 0, the simulated user on deepseek-v4-flash, and the programmatic evaluator only. The benchmark's natural-language assertions need a separate judge model, which we did not run, so rewards compare across our conditions and not to published leaderboards. Flash ran three trials per task, Haiku one.

IFEval [2] is 541 prompts each carrying one to three verifiable formatting or content instructions with a published strict checker. Here the checker is both the gate and the scorer, so this slice measures how repairable a failure is under a perfect gate, not the design of the gate. Haiku ran a fixed 250-prompt subset (seed 0).

Production replay is 74 real turns from one deployed product customer-service agent, June to August 2026. Every one of them is a failure: a person reviewing the draft before it was sent judged it wrong and edited it, and the failure was tagged as memory-shaped (a stale or missing fact). The set contains no turns the reviewer accepted as written, so its violation rates describe known failures, not production traffic, and the measure to read is how much of each failure the gates remove. Each turn is re-run through the same execution runtime with the deployed system prompt, the preceding thread, tools answered from that turn's recorded results, temperature 0, and a ~200-token record of the correct dated facts. These are the same 74 turns and the same baseline runs as the August replay study, re-scored.

2.2 Conditions

ConditionWhat happens to each agent output
No gatesRules live in the prompt or policy; the output goes straight out
Gates, specific feedbackOutput checked; on failure the agent receives which check failed, what was wrong and the valid values, and composes again; at most two repairs (one on production); the last attempt goes out
Gates, generic retryOutput checked; the feedback says only "Your last response failed a pre-send check. Try again." The repair limit and final-attempt release match the specific-feedback condition. The τ² generic-retry run uses the original ten gates, not the corrected nine-gate set.

The generic retry compares two repair setups. The comparison does not isolate the explanation itself: message role and whether the rejected draft was quoted also varied. On τ², the corrected nine-gate set was tested only with specific feedback. Comparing it with the ten-gate generic retry therefore changes both the gates and the feedback setup.

The original gates were derived from policy or instruction text, without reference answers, and fixed before the first runs. There were two later changes: a production leak gate was added after the first replay exposed the echo, and the τ² follow-up removed one gate and corrected another after inspecting the original results. The tables identify these conditions and retain the original rows.

2.3 Gates

All gates are deterministic and run against the exact candidate the pipeline is about to send. There are no LLM-graded gates in this study, for cost and so that the deterministic layer is measured on its own.

SliceGatesBasis
Productionfacts_consistent_with_context (every date, time, price and fee in the reply appears in the supplied facts, the tool results or the customer's own message) · no_unauthorised_commitment (pattern set for "an agent will call you" and its variants) · no_narration_leak (empty reply, internal notes, pipeline narration, echoed instruction text, duplicated body)The production judge design; the leak gate mirrors the production leak gate's intent and was added after the first run exposed the echo (section 4)
τ² retailauth_before_account_action · confirmation_before_write · order_status_precondition · write_once_per_order · cancel_reason_enum · item_lists_consistent · grounded_identifiers · no_fabricated_action · transfer_protocol · one_tool_call_at_a_timeEach maps to a quoted line of the retail policy
IFEvalThe prompt's own instruction checkers, strict modeThe published evaluator

2.4 Measures

The primary measure on every slice is the violation rate at send: the share of outputs (turns, conversations, prompts) that carry at least one violation after the pipeline has finished with them. On production this is a fabricated or mismatched value, an unauthorised commitment, or a leak; the fabrication scorer is independent of the gate, while the commitment and leak scorers are the gate's own checks, so on gated rows they measure what survived repair. On τ² it is every sent assistant message replayed through the nine harm gates with its real prior history, which for the prompt-only condition is independent of any gate. On IFEval it is the strict checker.

Task completion is τ²'s average reward and pass^k (the share of tasks solved in all k trials), and IFEval's loose accuracy. Costs are taken from the provider's per-generation accounting. Conditions are compared paired on identical turns or tasks, with 95% intervals from the paired differences. The August study put this production corpus's run-to-run spread at about ±10 points on a comparable metric; differences inside that band are treated as null.

2.5 Models

deepseek-v4-flash is the model under test on every slice and drives the τ² user simulator; claude-haiku-4.5 is the second model on τ² and IFEval. Grading is programmatic throughout. Hypotheses and metrics were written down before the first run.

3. Results

3.1 Production replay

Conditionfabricated valueunauthorised commitmentnarration leakany violationany, echo stripped
No gates6.8%45.9%17.6%56.8%56.8%
2 gates, specific feedback4.1%21.6%43.2%55.4%33.8%
2 gates, generic retry6.8%29.7%51.4%70.3%—
3 gates, specific feedback4.1%18.9%41.9%51.4%25.7%
3 gates, generic retry9.5%31.1%44.6%62.2%40.5%

n = 74 turns, deepseek-v4-flash. These were previously rejected drafts, so the rates do not describe overall production traffic. The two-gate rows ran before the leak gate existed. "Echo stripped" means the instruction-like prefix was removed before scoring; the experiment did not implement that fix in the sending pipeline.

Unmodified outputs: three gates with specific feedback reduced turns carrying any detected violation from 56.8% to 51.4%, a 5.4-percentage-point reduction. After removing the echo before scoring, the rate was 25.7%, a 31.1-point reduction (95% confidence interval −45.2 to −17.0), or 55% fewer than baseline. That larger result is conditional on echo removal. Neither measure establishes a reduction in human review workload.

Unauthorised commitments fell from 45.9% to 18.9% (−27 points, CI −40 to −14). Fabricated values fell from 6.8% to 4.1%, but the interval does not resolve that difference (CI −8 to +3).

The tested specific-feedback setup had fewer violations than the generic retry. The combined difference was 10.8 points on unmodified outputs and 14.9 points after echo removal (CI −27.8 to −2.0). Generic retry had more fabricated values than no gates, 9.5% against 6.8%; that comparison is descriptive. Section 4 explains the echo that repair failed to remove.

3.2 τ²-bench retail

Conditionrunsavg rewardsolved every trialsent a violationwhat got through
deepseek-v4-flash, 30 tasks × 3 trials
No gates900.83377%4.4%2 fabricated actions · 2 unconfirmed writes
Frozen 10 gates, generic retry900.78963%5.6%5 unconfirmed writes · 1 wrong order status
Frozen 10 gates, specific feedback900.72253%0.0%—
Precise 9 gates, specific feedback (post-hoc)900.82273%0.0%—
claude-haiku-4.5, 30 tasks × 1 trial
No gates300.63363%6.7%2 fabricated actions
Frozen 10 gates, specific feedback300.40040%3.3%1 fabricated action
Precise 9 gates, specific feedback (post-hoc)300.66767%0.0%—

A violation is a database write the customer never explicitly confirmed, a write on an order in the wrong state, an account action before authentication, or a message claiming an action was done when no write succeeded. "Solved every trial" is τ²'s pass^k. The intermediate run with only the unmeasured gate removed is described in section 3.5.

With the corrected nine-gate set and specific feedback, there were zero detected violations across 120 test conversations: 90 on flash and 30 on Haiku. This was a post-hoc follow-up on the same 30 tasks. Zero detected violations in this sample does not establish that the system will never violate a rule.

There was no statistically clear change in task completion: against no gates, average reward changed by −0.011 on flash (CI −0.060 to +0.038) and +0.033 on Haiku (CI −0.14 to +0.21, one trial). Those intervals allow meaningful decreases as well as increases; they do not establish equivalence. No gates was inconsistent across its three trials on 3 of 30 tasks, compared with 4 of 30 for the corrected gates.

Measured flash agent cost per conversation was similar: $0.0025 with corrected gates against $0.0024 without, with 24.9 against 25.4 messages. Section 3.6 reports costs across the study.

With the original ten gates, the generic retry sent detected violations in 5.6% of flash conversations, compared with 4.4% without gates and 0.0% with the same ten gates and specific feedback. In some failed retries, the model proposed another write without asking the customer. These results favour the tested specific-feedback setup, but do not isolate the explanation as the cause.

3.3 IFEval

Modelno gatesspecific feedbackgeneric retryfailures left
deepseek-v4-flash88.9%96.3%94.3%11.1 → 3.7 → 5.7%
claude-haiku-4.585.6%92.8%92.4%14.4 → 7.2 → 7.6%

Prompt-level strict accuracy on 541 prompts (flash) and 250 (Haiku); 95% intervals are ±2–3 points on flash and ±4 on Haiku. "Failures left" reads no gates → specific feedback → generic retry. The gated runs use 1.4× (flash) and 1.2× (Haiku) the output tokens of no gates.

Gates with specific feedback removed two thirds of the residual instruction failures on flash and half on Haiku, and loose accuracy, the completion proxy, rose with them. Of the 59 flash prompts that failed first time, the first repair fixed 30 and the second a further 9; on Haiku, 14 and 3 of 35. The specific-feedback setup scored 2.0 points above generic retry on flash and 0.4 on Haiku. These differences do not establish a clear advantage. One possible explanation is that IFEval already puts the instruction in the prompt. The failures that survived two repairs included letter frequency, capitalised-word counts and exact word counts.

3.4 Statistical comparisons

Paired differences whose 95% interval excludes zero:

ComparisonΔ95% CI
Production, any violation (echo stripped): 3 gates with specific feedback − no gates−31.1 pts−45.2 … −17.0
Production, unauthorised commitment: 3 gates with specific feedback − no gates−27.0 pts−40.3 … −13.8
Production, any violation (echo stripped): specific feedback − generic retry−14.9 pts−27.8 … −2.0
Production, unauthorised commitment: specific feedback − generic retry−12.2 pts−24.1 … −0.3
Production, fabricated value: specific feedback − generic retry−5.4 pts−10.6 … −0.2
τ² flash, avg reward: frozen 10 gates − no gates−0.111−0.217 … −0.006
τ² Haiku, avg reward: frozen 10 gates − no gates−0.233−0.414 … −0.053
τ² flash, avg reward: 9 gates with the imprecise confirmation gate − no gates−0.067−0.124 … −0.009
IFEval flash, strict accuracy: specific feedback − no gates+7.4 pts≈ +4 … +11
IFEval Haiku, strict accuracy: specific feedback − no gates+7.2 pts≈ +2 … +12

Differences that remain unresolved include production fabrication, three gates with specific feedback against no gates (−2.7 points, CI −8 to +3); τ² completion with the corrected gates (−0.011 flash, +0.033 Haiku, both intervals crossing zero); and IFEval specific feedback against generic retry on flash (+2.0 points). An interval crossing zero means these tests do not resolve the direction of the difference; it does not prove equal performance.

3.5 What the bad gates cost

The frozen τ² set carried two defects. Both were found in the numbers, not in the code.

The first gate measured nothing. "You should at most make one tool call at a time" is a line in the retail policy, and the reward never checks it. Haiku batches its lookups in parallel, so the gate fired on 28 of 30 tasks, and the frozen condition lost eight tasks and won one against prompt-only. Each block moved a conversation off a path that would otherwise have completed. Dropping that gate alone recovered most of the loss on Haiku, 0.400 to 0.600, and some on flash.

The second gate fired when the harm was absent. The frozen confirmation gate's refusal pattern matched a bare "no" anywhere in the customer's last message. The policy's own accepted cancellation reason is "no longer needed", and customers write "no other items to change". With the first gate removed, all 26 of this gate's firings on flash were false positives, and before rescoring it had put prompt-only's apparent violation rate at 16.7% rather than 4.4%. Rewriting the pattern so that a bare "no" counts only as a standalone refusal, the precise set in the tables, recovered completion to estimates close to or above baseline, with wide confidence intervals. The prediction for that run, no violations and reward within noise, was written before it ran.

Together the two gates cost up to 23 points of task completion on Haiku and prevented no violation.

3.6 Cost

The whole study cost $17.25: production $0.75, IFEval $3.30, τ² $11.89, build and smoke runs $1.31. Per unit, the gated pipeline uses between 1.0× (τ² flash) and 1.4× (IFEval flash) the output tokens of prompt-only. The difference is the repair rounds, which fire on 10 to 14% of units.

4. Negative and null results

Reprompting did not fix the echo. In the production replay, deepseek-v4-flash sometimes prefixed its reply with " Never guess." or " Do not guess or invent." The leak gate detected the prefix, but the model repeated it in 30 of 31 repaired turns. A second repair was not tested on this slice. Removing the prefix in code is a candidate fix; this experiment removed it only before scoring, not in the sending pipeline. Both the unmodified and echo-stripped results are reported in section 3.1. In deployment, unresolved failures would need the configured hold or suppression behaviour.

No consistency gain. We predicted that gates would make repeated runs more repeatable, with a wider gap on pass^k than on pass^1. On τ² flash, pass^3 is 0.73 with precise gates and 0.77 without, and per-task inconsistency is 4 of 30 against 3 of 30. Null at this n.

The weaker model did not gain more. Haiku gained 7.2 points on IFEval, flash 7.4. Null.

Generic retry left violations unresolved. The tested generic retry had more detected violations than the ungated baseline on τ² flash and more fabricated values on production replay. The study does not establish that every generic retry is worse, or that the explanation alone accounts for the difference. A retry limit also needs a clear policy for unresolved failures.

An instrument failure. The first Haiku IFEval pass delivered the repair prompt as a trailing system-role message, and every repair came back empty, at 3 output tokens. All Haiku repairs were re-run with the prompt as a user-role message prefixed [PRE-SEND CHECK]. The system-role outputs are kept on disk and excluded.

5. Implications for deployment

  1. Test checks against the harms they are meant to prevent. The τ² results show how an unnecessary restriction or a false match can reduce completion. A benchmark score alone does not establish whether a rule is needed in deployment.
  2. Measure false positives before shipping. Run each check over outputs a person has approved, inspect incorrect blocks, and revise the check before relying on it.
  3. Choose a fix the system can perform. Test simple code fixes for failures such as an echoed prefix. Use model feedback when the answer itself needs to change. Echo removal still needs validation in the sending pipeline.
  4. Give specific feedback and limit retries. Identify the failed rule and what would satisfy it. This setup performed better in the reported comparisons, although the explanation's separate contribution remains unresolved. We tested at most two repairs, and only one on production replay.
  5. Define what happens after the last failed check. Hold for human review or suppress according to the rule. The experiments released the final attempt to measure residual violations; they did not measure review-queue load.
  6. Run inexpensive checks first. Our production design puts string and tool-call checks before model-graded rules. The added cost and benefit of those model-graded rules were not measured here.

6. Limitations

  • Budget. This study ran on a $30 cap, which set the slice sizes and the model panel: 74 production turns, 30 τ² tasks, 541 and 250 IFEval prompts, one seed. With more budget the same harness extends to the full τ² retail and airline sets, more trials and a wider panel; that is the difference between a directional result and a resolved one.
  • Statistical power. The earlier production replay found run-to-run spread of about ±10 points on a comparable metric. That is separate from the paired confidence intervals reported here. τ² Haiku ran a single trial and has a wide completion interval. We did not test a predefined equivalence margin, so a statistically unresolved change is not evidence that completion is unchanged.
  • Post-hoc gate changes. The corrected τ² set was defined after the original results and rerun on the same tasks. Its prediction was written before the follow-up, but it remains a post-hoc result that needs validation on new tasks. The production leak gate was also added after the first replay exposed the echo.
  • Feedback comparison. Explanation, message role and quoting the rejected draft were confounded. The corrected τ² gate set also lacks a matching generic-retry run. The results compare tested setups, not the isolated effect of the explanation.
  • Echo removal and workload. The 55% relative reduction uses outputs with the echo removed before scoring. Unmodified outputs fell from 56.8% to 51.4%. Automatic removal in the sending pipeline and changes in human review workload were not tested.
  • Scorer overlap. On production, the commitment and leak residuals are scored by the gate's own checks; only the fabrication scorer is independent. On IFEval the gate is the scorer by construction. On τ², the violation measure for the prompt-only baseline is independent of any gate, and the task reward is independent on every condition.
  • τ² evaluator. Programmatic reward only; rewards are not comparable to Sierra's published figures, which include natural-language assertions graded by a judge.
  • Model panel. Two cheap models; no frontier model was tested.
  • One repair on production. The replay runner allowed one repair where the spec allowed two. The effect of an additional repair on this slice was not measured.
  • Deterministic gates only. The production system also runs LLM-rubric rules after the string and tool-call layer. Their contribution, cost and false-positive rate were not measured here.

7. Further research

Our production pipeline already runs the shape studied here, string and tool-call checks first and LLM rubrics last, and its configuration surface, per-rule consequences, one rule vocabulary shared between tests and live gates, and a per-attempt audit trail, is the instrument for the next experiments. Each item below is a gap this study exposed and did not close.

  • False-positive rate as a shipping gate. Nothing in the production system measures a check's false-positive rate; test mode holds every reply for a human but never reconciles the human's decision with the rules' verdict. The August precision gate for mined rules (fail at least 3 drafts, pass at least 90% of human versions, pass at least 95% of approved controls) is the instrument. The experiment is to apply it to every existing gate, including the leak patterns and the two style gates, against the approved-send corpus, and report which would have been rejected.
  • Auto-fix versus reprompt, per gate. Classify each deterministic gate by whether its failure is a string edit (echo prefix, dash replacement, markdown separator, whitespace) or a different answer, implement the strip for the first class, and measure the change in repair rounds, approval-queue load and violation rate on the replay corpus. Test whether the fix reduces repair rounds and residual violations without altering valid replies.
  • LLM-rubric rules: cost, precision and placement. Rubric rules are graded by a model on every rewrite round. Measure their false-positive rate on approved sends, their marginal catch over the string and tool-call layer, and the effect of running them last and only on turns that pass every deterministic rule. The August rubric-rule result (first-pass check rate 39%, 249 repairs) suggests they are the dominant source of repair churn.
  • The repair prompt itself. Three variables were confounded here: reason against no reason, system role against user role, and whether the rejected draft is quoted back. The production system quotes the draft and uses different roles on different executors. A 2×2×2 on the replay corpus would settle which components carry the effect, and whether the system-role empty-reply behaviour generalises beyond one provider.
  • Consistency at power. H2 was null at 30 tasks × 3 trials. Repeating the τ² slice at 10 trials, and adding the airline domain, would resolve whether gates narrow run-to-run variance.
  • A gate for new gates. Both τ² defects were invisible in code review and obvious in the numbers. Specify a minimal bench that any new gate must pass before it enters the auto-send list: fires on a seeded violation set, does not fire on the approved-send corpus, and has an auto-fix or a repair hint attached.
  • Corrections into gates. Close the loop the design implies: a human rejection or edit proposes a check, the precision gate vets it, and an accepted check lands in both the agent's test suite and its live gate set. Measure yield per hundred corrections and the violation rate over time.
  • Frontier models and the tool-call convention. Haiku's parallel tool calls are the reason the unmeasured gate hurt it. Re-run τ² with a frontier model, and treat "one call at a time" as a property of the executor rather than a gate.

References

  1. Barres, V., Dong, H., Ray, S., Si, X., Narasimhan, K. τ²-Bench: Evaluating Conversational Agents in a Dual-Control Environment. 2025. arXiv:2506.07982.
  2. Zhou, J., Lu, T., Mishra, S., Brahma, S., Basu, S., Luan, Y., Zhou, D., Hou, L. Instruction-Following Evaluation for Large Language Models. 2023. arXiv:2311.07911.
  3. Yao, S., Shinn, N., Razavi, P., Narasimhan, K. τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. 2024. arXiv:2406.12045.
  4. Hivemind. Memory over context: accurate, low-cost long-term recall for LLM agents. 2026. www.hivemind.au/insights/memory-over-context. The August replay study and precision gate referenced in sections 2.1 and 7.

Appendix: research summary

We tested deterministic output checks with specific feedback on τ²-bench retail, IFEval and 74 previously rejected production turns, using two inexpensive models. The corrected nine-gate τ² setup had zero detected violations across 120 conversations, with no statistically clear change in completion and similar measured flash agent cost. The corrections were made after the original results and tested on the same tasks.

On production replay, three gates reduced turns with any detected violation from 56.8% to 51.4%. Removing an instruction-like echo before scoring reduced the rate to 25.7%, or 55% fewer than baseline. That automatic fix was not implemented in the sending pipeline, and human review workload was not measured.

The tested specific-feedback setup performed better than generic retry in several comparisons. Message role, quoting the rejected draft and, in the corrected τ² comparison, the gate set also varied. The explanation's separate contribution remains unresolved. Poorly designed gates reduced task completion; repeated feedback did not reliably remove the echo.

Appendix: measurement techniques

Every run appends to an account-level spend ledger with a hard stop at the budget cap, so no slice can exceed it silently. τ² gates are replayed offline over every saved conversation with its real prior history, so the violation measure for any condition can be recomputed without model calls, and was, after each gate correction. The IFEval repair message quotes the same counts the checker computes, mirrored expression by expression, so the model is never told a different target from the one it is graded on. All 25 IFEval instruction types were exercised offline against all 541 prompts before any paid call. Hypotheses, conditions and metrics were fixed in SPEC-gates.md before the first run; every correction, including the two wrong τ² baseline figures that preceded the confirmation-gate fix, is logged with its date in DECISIONS.md.

Running agents in front of customers?

This is the check layer we build into the systems we run for clients. If your agents send messages or take actions on a customer’s behalf, we can show you what the gates look like on your own rules.