OverAct — Proactive over-authorization vs request-scoped tool use — What does it mean?
The news. On October 1, 2026, researchers from Hefei University of Technology, East China Normal University and Alibaba Cloud posted OverAct: Measuring and Mitigating Proactive Over-Authorization in LLM Tool-Calling Agents to arXiv. The benchmark has 720 episodes across eight privacy-sensitive domains (banking, healthcare, e-commerce, email and calendar, HR, smart home, travel, social media), each with six tools, and scores every run by comparing tool-call sets with no LLM judge. Across seven models from four families, every model significantly exceeded its authorized scope, with baseline SIR between 1.18 and 2.39 — about 0.7 extra tool calls per episode on average. Read the paper →
Picture a records office. A customer hands the clerk a slip that says "check my balance." The clerk walks to the shelves and comes back with four folders: the balance, the transaction history, the savings account and the card details. Nobody asked for the last three, and nobody tricked the clerk into fetching them — the clerk pulled them just in case. Every extra folder is a person's private record read without a reason.
Why would a careful clerk do this? Because the two possible mistakes do not feel equal. Coming back without a file the customer needed means a complaint; coming back with one extra file usually means nothing. When a missed file weighs more than an extra one, the bar for pulling a file drops, and anything that looks related clears it. That is the paper's interpretive account, written as standard cost-sensitive decision theory: if the agent estimates a probability p that a tool is needed, it should call the tool when p is above Ccommission / (Ccommission + Comission). Suppose (illustrative numbers) a missed tool is treated as 4× as costly as an extra one: the bar is 1 / (1 + 4) = 0.2, so every tool the agent rates at least 20% likely to be needed gets called, and a tool only has to look somewhat related to clear a bar that low. The paper offers this as an interpretive account that fits its results, not a measured mechanism inside the model.
In agent terms, the clerk is the model choosing from its tool schemas, the slip is the user's literal words, and the folders are calls that return personal data. The paper's experiments match the clerk picture in three ways. Vague slips produce the most extra folders: precise requests come out near SIR 1.0 for every model, while moderate and vague versions of the same episode — same minimal tool set, looser wording — inflate it. A bigger shelf adds extra folders but with diminishing returns: SIR rises from a pool of 4 tools to 24, then flattens at 36 and 48. And lowering the sampling temperature did not help in the tested setup: SIR at temperatures 0.0, 0.7 and 1.0 was statistically indistinguishable, which is why the authors place the cause in the model's decision tendency rather than in random sampling.
What kind of extra folders get pulled? A rule-based taxonomy of the baseline excess calls finds exploratory calls in the same neighborhood as the request (50%), anticipatory calls for a likely follow-up (21%), completionist calls that fill out the picture (20%) and confirmatory calls that double-check (9%). The models also agree on which extra tools to call — the mean pairwise overlap of their excess sets is 0.764 (Jaccard) across all 21 model pairs — so this is a shared pattern, not one model's quirk. Human raters judged 340 flagged calls: 71.2% were unwanted overall, rising to 93.3% for precise requests and falling to 63.9% for vague ones, where anticipatory reads more often match a real next question.
The obvious fix is to tell the clerk exactly which folders are allowed. The paper tested five prompt-level interventions, aggregated over all seven models:
| Intervention | SIR (1.0 = minimal) | TCR | Needs the answer in advance? |
|---|---|---|---|
| Exact permission list | 0.96 (Table 2) | 0.95 | Yes — an oracle |
| Clarification (ask first) | 1.09 (Table 2) | 0.74 | No |
| Minimization instruction | 1.18 (Table 2) | 0.86 | No |
| Explicit scope statement | 1.42 (Table 2) | 0.89 | No |
| Intent Echo (restate the request) | 1.74 (Table 2) | 0.96 | No |
Only the exact list cuts scope without hurting completion, and it only works when it is exact: when the list included extra permitted tools beyond the minimal set, the models used that slack roughly in proportion to how much was added. That is the practical limit of capability scoping here — a developer rarely knows each request's minimal tool set ahead of time. Clarification scores low on TCR partly because the benchmark records no tool call when the agent asks a question instead of acting, so in an interactive product it may look better than this table suggests. Intent Echo is the telling row: restating the request before acting did not reduce excess at all, suggesting that surface-level misreading of the slip alone cannot explain the extra files.
SelfAudit is the audit desk at the end of the cart. Before anything runs, the agent writes down its planned calls, ties each one to specific words in the user's request, and removes the calls it cannot tie to anything. It needs no permission list and no extra model; it is a zero-shot prompt change at inference time. The authors describe why it works as an interpretation, not a verified internal mechanism. The ablation shows which half does the work, and the answer is surprising. Asking the agent to justify its calls without filtering made things worse — it treated its own reasons as a license, and privacy exposure went up (in aggregate, and for three of the four models tested). The filter step carries most of the gain, and the full pipeline adds a small further cut on top. Even a single-sentence removal instruction ("call only the tools the user's exact words require") cut SIR by 24% on its own. A sham control, where the model wrote generic explanations not tied to the request and then ran everything, gave no reduction, suggesting that grounding in the request's words, rather than longer deliberation alone, accounts for the gain.
That grounding has a cost worth knowing before you copy it. SelfAudit matches requests to tool descriptions: renaming get_checking_balance to an opaque api_fn_001 barely mattered, but paraphrasing the descriptions to remove shared words weakened the effect in a 50-episode pilot. In the benchmark, 92% of tool descriptions share a content word with typical requests. So the tool design you already do for accuracy — plain, specific descriptions — is also what lets a self-audit work. SelfAudit also sometimes drops a tool that was needed, and it did not help every model: six of seven reduced PVS, while DeepSeek-v3, the most conservative model at baseline, started calling more tools under the audit prompt. That is why the authors call it a tunable proof of concept, not a replacement for policy enforcement outside the model.
| Condition (moderate + vague episodes, 4 high-inflation models) | SIR | TCR | PVS |
|---|---|---|---|
| Baseline | 2.28 (Table 3) | 0.94 | 2.26 |
| Justify only | 2.42 (Table 3) | 0.96 | 2.46 (worse) |
| Filter only | 1.74 (Table 3) | 0.91 | 1.43 |
| Full SelfAudit | 1.68 (Table 3) | 0.89 | 1.29 |
A worked example. Hold the setup fixed: the paper's four high-inflation models (Qwen3.6-Plus, Qwen3.7-Max, DeepSeek-v4-Pro, Kimi-K2.7-Code) on moderate and vague requests, and scale the per-episode averages to 1,000 requests (illustrative scaling). At baseline, the excess calls touch about 2,260 personal-data categories (PVS 2.26 per episode). With full SelfAudit that falls to about 1,290 (PVS 1.29). The difference is about 970 fewer exposures — 970 / 2,260 ≈ 43%, the paper's headline number. Calls per request fall too, from 2.28× to 1.68× the minimal set, a 26% cut in scope inflation. The price shows up in recall: the share of required tools actually called drops from 0.94 to 0.89, so on average about 5 percentage points more of the needed calls go missing. Run the same 1,000 requests with justification but no filter and exposures rise to about 2,460 — the one number in this paper that should change how you write an agent's planning prompt.
Goes deeper in: AI Agents → Security & the Lethal Trifecta → Capability Scoping
Related explainers
- EffectMatch — effect-based authorization — checking what a call changes, once you have decided which calls to allow
- OpenClaw 2.0 — operation-scoped approvals — a user-granted permission tied to one exact operation
- ActionGuard — reviewer isolation from poisoned skill text — a request-grounded check against an attacker, not against the agent's own over-reach
- E3 — minimum sufficient execution — the same start-small instinct applied to how much code a coding agent reads