DelegationBench — Matched-pair act-or-ask testing — What does it mean?
The news. On October 4, 2026, Shiva Pochampally posted DelegationBench: Measuring When AI Agents Should Ask Before Acting on arXiv. It has 156 scenarios across files, calendars, messaging, purchasing, publishing and code; 96 of them form 48 matched pairs, 12 for each of four features: scope (did the user ask for this action), stakes, reversibility and visibility to other people. Ten models from five families were run five times on every scenario, for 7,800 primary evaluations; a tool-sandbox follow-up presented 48 of the scenarios as tasks using simulated tools. Read the paper →
Picture a new house-sitter who finds a note: "clear the garage." One version says the junk goes into the recycling bin, where it can be pulled back out. The twin note is word-for-word the same, except the junk goes on a bonfire. A good sitter goes ahead with the first note and phones the owner about the second — and handing them both notes, then checking whether the answer moves, is how you test that they react to the detail that matters. That is a matched pair, and it is the core of DelegationBench: the paper's own example is moving a file to the Trash, where it can be restored for 30 days, versus deleting it permanently under the same instruction.
Now picture a lazier test. You give the sitter fifty different notes, write down what a panel of neighbours would have done with each one, and score the sitter on how often they match. A sitter who phones home whenever the note contains a scary word like "burn" can score well on that test without ever reading what the note actually asks. That is the agreement score, the standard way to evaluate an agent's policy enforcement: show a proposed action, ask whether to proceed, compare with human labels.
The paper also splits the sitter's "phone the owner" into two different calls. Ask means "may I do this?" — the action is clear and the agent needs permission. Request means "which box did you mean?" — the agent needs a missing fact. Merging them hides which of the two a model is missing, and the two need different fixes in a harness.
Three ways the agreement score misleads
Responsiveness gap. The author wrote a three-line keyword rule after looking at the benchmark: choose Request if one of 15 missing-information phrases appears, otherwise Ask if one of 14 risk words appears outside the instruction, otherwise Act. It is a deliberate counterexample, not a baseline. It agrees with the annotators on 69.1% of scenarios — more than eight of the ten models — yet its decision changes in only 9 of the 48 matched pairs. The same pattern shows up among real models: Gemini 3.5 Flash-Lite has the highest agreement (73.4%), but its act rate does not change on any of the 12 stakes pairs. Across the ten models, agreement and pair responsiveness are negatively correlated (Spearman ρ = −0.529, a rank correlation where −1 means the two orderings are exact opposites; n = 10, descriptive only).
Elicitation gap. The same 48 decisions, asked in five equivalent formats, move a single model's act rate by between 4.2 points (Claude Opus 4.8) and 52.5 points (GPT-OSS-20B). Shuffling the order of the four options gave the closest match to the annotators' vote shares for nine of the ten models, so the paper recommends counterbalancing option order and reporting the spread across formats.
Judgment-to-action gap. Every model stopped to ask the user less often when it had to do the task with tools than when it judged a shown action. Gemini 3.5 Flash-Lite asked in 47.5% of judgments but in 4.2% of Tool-Loop episodes, even though every episode offered a structured way to ask. The settings differ in several ways at once (the proposed action is hidden, the format changes, permission is never mentioned), so the paper says judgment results do not carry over to action, not which difference causes the drop. This is the same reason a production team runs a new policy in shadow mode against live traffic before trusting an offline score.
| Gap | What a single agreement score hides | Largest effect reported | What to report instead |
|---|---|---|---|
| Responsiveness | Whether the decision reacts to the feature that should matter | Keyword rule: 69.1% agreement, 9 of 48 pairs moved (§5.1) | Share of matched pairs moved, per feature |
| Elicitation | Whether the decision survives a rewording of the same question | Act rate moved by up to 52.5 points across formats (§5.2) | Range across formats, with option order counterbalanced |
| Judgment-to-action | Whether the model asks when it actually holds the tools | Ask rate 47.5% when judging vs 4.2% with tools, one model (§5.3) | Ask rate in a tool loop, measured separately |
The worked example: 9 pairs out of 48
Hold three things fixed: the same 48 matched pairs, the same annotator labels, and the same agreement metric. The keyword rule moves its answer in 9 of 48 pairs, about 19%. The least responsive model moves in 45.8% of pairs, which is 22 of 48; the most responsive moves in 70.8%, which is 34 of 48. So the rule gives both twins the same answer in 39 pairs, while even the least responsive model fails to move the designed way in only 26 — and the rule still outscores eight of the ten models on agreement, 69.1% against a model range of 50.6% to 73.4%. For scale, a policy that always answers Act gets 46.3% and one that always answers Ask gets 47.7%. A leaderboard sorted by agreement would rank the rule near the top; a leaderboard sorted by pair responsiveness would put it last. That is the whole argument for building golden cases as twins rather than as a loose pile.
What it does not show
The gaps are not a general failure to follow rules. When the correct answer follows from a written policy, the three models tested follow it almost perfectly (97.3% to 100% balanced accuracy, which weights accuracy equally on cases where acting is permitted and cases where approval is required). And when a Tool-Loop episode ended with a simple permission question, a plain "yes, go ahead" led to the tool being called in 117 of 120 cases (97.5%). The trouble is the ambiguous middle — no written rule, a judgment call — which the paper argues is the common case in practice. The benchmark has limits too: 156 hand-written scenarios, three student annotators who themselves agree only moderately (Krippendorff's α, a chance-corrected agreement measure where 1 is perfect, is 0.437 across all four labels), each feature built from just three scenario families, and 13 of the 24 stakes and reversibility pairs changing something besides the target feature. The paper treats the annotators as an informed reference point, not ground truth. A closely related agent skill — deciding when to stop entirely rather than when to ask — is covered in agentic abstention.
Goes deeper in: Agent Engineering → Layered Guardrails → Policy Enforcement