An open 27B model, not fine-tuned, goes from 49.5 to 57.2 pass^1 on τ³-bench banking_knowledge. That is above the top public entry, 55.2. The model, its weights, the tools and the benchmark's policy documents are unchanged; only the layer is added.
- +7.7
- Performance. pass^1 on banking from the layer alone. Tasks failed in all four trials: 35 → 25.
- 1 file
- Cost to move. A new domain is a new declaration file. The engine holds no domain words and stays the same.
- 0 blocks
- Reliability. correct actions wrongly blocked on 20 unseen airline tasks. Every decision is deterministic and logged.
The failure
Slips the agent knows better than
A customer-service agent reads the policy, calls tools, and writes to a database. Most of its failures are not missing knowledge. They are small, checkable slips: a write before the required lookup, a sum done in the model's head, a value that nobody said. One slip is enough, because a write to a record cannot be taken back.
get_reservation_detailscreated_at 2024-05-02 · policy clock 05-15cancel_reservation(…) takes the customer's wordreservation cancelled
DENY · NOT EXECUTED
✕ booked within 24 h?
record 05-02 · policy clock 05-15
① the unexecuted call cancel_reservation
② the fact from the record created_at 2024-05-02
③ the policy's cancellation sentence, verbatim
record unchanged
Recorded refusal · the same rule on τ² airline task 43 (excerpt)
…The reservation record read in this conversation shows created_at 2024-05-04T07:38:29 (current time stated in the policy: 2024-05-15 15:00:00), cabin basic_economy, insurance no; flight status reads in this conversation showing a segment of it cancelled: 0. …
The only values in it are facts read from records in this conversation. The policy quotation is omitted from this excerpt.
VOCAB_EXT.md §5-2.The idea
Enforce only what a program can decide
Some policy conditions can be settled by looking: did a lookup happen, what a field in a fetched record says, whether a value appears in a tool output. Others need judgment: is this request reasonable, what did the customer mean. The layer enforces the first kind and leaves the second kind to the model.
CLOSEDthe layer enforces
The answer is fixed by events, record fields and strings in the transcript.
- eventWas the reservation read before
cancel? - recordIs
created_atwithin 24 h of the policy's clock? - stringDoes the ZIP code being written appear in a tool output or customer message?
- arithmeticDoes the refund equal the value computed from the fetched records?
OPENthe model decides
The answer needs interpretation. The layer may show the rule text, but never blocks on it.
- intentDoes the customer really want to close the account?
- toneIs this a complaint that should be escalated?
- relevanceWhich of three products fits what they described?
- meaningIs "a few days ago" inside the window?
The mechanism
One turn, step by step
As you scroll, the part of the figure for each step lights up and a dot follows the path the information takes.
The model reads the conversation and proposes its next turn: a tool call or a message. Nothing has run yet.
The engine loads the domain's declaration file and the transcript so far: which tools ran, what they returned, what the customer said.
Seven levers test closed conditions on the proposed call. Each lever either stays silent or reports one violated fact. Each lever is explained in "Seven levers" below.
No violation: the call goes to the real tool unchanged. The layer never edits arguments.
A violation: the call does not run. The model receives the call, the unconfirmed fact and the policy sentence, and proposes again. A small fixed budget caps retries, so the dialog never stalls.
Every pass, denial, injected computation and retry is written to an audit log with the tool output it relied on. The same input gives the same decision.
Why the refusal carries no value
If the layer wrote the fix ("cancel with refund 0"), a mistake in a rule would become a mistake in an action. Returning only the fact and the policy sentence keeps the model responsible for the next move, and keeps the layer from acting as a second agent.
Seven levers
One lever per kind of slip
The levers came from reading failed trajectories and naming each recurring slip. Pick one to see what it reads, what it decides, and what happens to the score when it is switched off.
ABLATION*.md.Moving to a new domain
The engine stays. One file changes.
The engine contains no domain words: no "card", no "flight", no "fee". Everything specific to a domain lives in one declaration file: which lookups a write needs, which values must be sourced, which computations are available, which policy sentences go with which tool.
hours_between · count_of · prefix · eq) and one operand (now) added to the engine (+74 lines). After the change, banking decisions over 6,197 recorded turns were byte-identical. Source: VOCAB_EXT.md.What the numbers say
Gains where models break closed rules
We swapped only the declaration and ran the same layer on two more benchmarks and ten model configurations. The gain is large where the model commits closed-rule slips and disappears where it does not.
Unseen airline tasks
Rules written from 30 training tasks only, scored on 20 unseen tasks. GPT-4o-mini 31.2 → 52.5, with 0 correct actions blocked.
Against a published gate, same harness
Reason Less, Verify More (2607.07405) ships four airline-specific gates. Re-run under our simulator and tasks. Simulations lost on fired tasks: this layer 1, the gates 5.
Where it does not help
Three places with no gain
- ≈0
- τ² retail, all 4 models. Failures there need judgment (which item, which variant) or come from the user simulator ending early. Few are closed-rule slips.
- ≈0
- τ² airline, Qwen3.8-27B. The model already follows the policy. The layer fired on one task, net zero. A strong model leaves little to catch.
- −0.4
- SOPBench hotel. +11.0 under the official scorer, −0.4 after fixing an OR→AND scorer defect on 56 tasks. The official gain is an artefact.
Where the layer does not help it does not hurt: losses of 0 to 3 simulations out of hundreds, none traced to a wrong denial.
How we measured