Agent·

PACE — Influence-path confinement vs content vetting — What does it mean?

The news. On October 1, 2026, researchers from KAUST, RIKEN, the University of Macau and other institutions posted PACE: Provenance-Aware Capability Enforcement for Tool-Using LLM Agents to arXiv. PACE mediates each tool call in four phases — propose, cut and certify, enforce, finalize — and was tested on eight agent-security benchmarks (including AgentDojo, InjecAgent, WASP, PASB and MCPTox) with three model families. The evaluated configuration had the strictly lowest attack success in 62 of 79 eligible attack columns and tied in 14, while full-benchmark task utility fell by at most 2.92 points against an undefended agent. Read the paper →

Picture a building's water system. One tank holds water from an untrusted supplier; another is the private reservoir. Pipes run from both into a mixing room and out to a tap on the street. Testing the supplier's water once, when the contract is signed, settles less than it seems: two batches can pass the same test and still behave differently once they are in the pipes. The paper makes this precise for agents. A skill or tool description is prose the agent will later follow, so a reviewer that reads it is reading the attacker's own words, and a safe variant and a leaking variant can produce the same review evidence. Under two stated conditions, the authors prove that a sound gate then cannot relax its checks for either variant.

So the plumber works at the last point that still matters: the tap, right before water leaves the building. There the plumber has the pipe map, knows which tanks are dirty or private, and can still close a valve. In agent terms, that is the moment before a tool runs, when the frozen arguments, the data already read and the user's authenticated request all exist together — none of which is available when a skill is first installed. It is the data-flow graph from the security module, built automatically and checked on every call.

Influence-path confinement asks one structural question: is there any path in the pipe map from a dirty or private tank to this tap? PACE builds that map as a provenance graph. When a call is proposed, each argument is frozen and matched against values already seen in the run; a match links it to the tool return, file or memory record it came from. Anything an artifact, tool return, retrieved page or memory record supplied is marked untrusted, and anything derived from it inherits the mark. Secrets and personal data form a second source set, and arguments are also scanned for registered secrets under a fixed family of re-encodings, so a lightly encoded key is still caught. The authors state the limits plainly: paraphrase, translation and splitting a value into fragments fall outside that scan, and the graph only holds paths the host and schemas represent — a value the model rewrites in its own words can lose its link to the source.

Then comes the plumbing step. Old pipes cannot be rerouted: edges from past calls are uncuttable, because the executor cannot undo a call that already ran. Only the new edges this call would add are candidates, and each one counts as a valve only if the deployment declares an action that really severs it. PACE computes the minimum cut — the cheapest set of those valves that disconnects every source from the unsafe outlet — and when no argument edge can be severed, the cut falls back to blocking the call itself. A call built only from the user's request and the host has no path to cut, though it still faces the effect check below. This is the cut-a-leg idea applied per call rather than designed once.

A path alone is not a verdict: reachability says an effect may occur, not whether the user asked for it. So the plumber also reads the work order at the tap. Each call is expanded through its schema into effect atoms and compared against a capability compiled only from the authenticated request, never from model output. Nine obligations are checked, including authorization coverage, confidentiality, persistence, delegation attenuation, budget and freshness. This half is close to effect-based authorization, and it is the policy enforcement layer of a guardrail stack, driven by provenance rather than by a fixed rule list.

The paper's own ablation is candid about which half does the work. On a 1,167-case subset with the local Qwen-3.8-27B model, the effect check carries most of the security gain, and the full combination shows no security benefit beyond the best reduced arms — its value is utility. Path confinement alone also drives AgentDojo attacks to zero, at the cost of most legitimate tasks; the effect check restores calls the path check blocked when the request authorizes them. In practice the cut rarely rewrote an argument: on AgentDyn, 5,842 non-empty candidate cuts across 6,647 guarded calls produced no executed argument-level rewrite, so the path check proposed whole-call blocks rather than argument rewrites, and the effect check could then restore the calls the request authorized. A third layer, boundary adaptation, adjusts these decisions at the point of execution; adding it to the effect check cut the MCPTox refusal rate from 99.3% to 1.6% with no valid attack succeeding in either arm.

One distinction matters when reading the guarantees. The paper defines a certified contract, in which the final call must keep the cut and pass the effect check again, and proves path separation and effect soundness for it — only over the paths the schemas represent, and explicitly not full noninterference. Every experiment instead runs PACE-p, which may restore a blocked call or apply a declared repair without rechecking the final call, so the measured numbers and the proofs describe slightly different systems.

Ablation arm (Qwen-3.8-27B subset)AgentDojo ASR ↓AgentDojo utility ↑PASB ASR ↓
No defense (Table 2)37.4%76.8%41.4%
Path confinement only (Table 2)0.0%13.6%31.0%
Effect check only (Table 2)0.0%64.6%0.0%
Path + effect (Table 2)0.0%65.2%0.0%
Full PACE-p, as evaluated (Table 2)0.0%67.2%0.0%

A worked example. Hold three things fixed: the local Qwen-3.8-27B model, the paper's 1,167-case ablation subset, and the AgentDojo rows above. Scale the rates to 1,000 attack runs and 1,000 benign tasks (illustrative scaling; attack and utility rates come from different cases). With no defense, about 374 attacks land and 768 tasks finish. Path confinement alone lets no observed attack through in this subset but finishes only 136 tasks — 632 fewer. Full PACE-p also lets no observed attack through and finishes 672 tasks. Net: about 374 attacks stopped for about 96 lost tasks, against 632 lost tasks for the path check alone. Zero observed successes on a fixed test set is not a guarantee against future attacks. The effect check is what turns a blunt blocker into a usable gate. On the full benchmarks the loss is smaller (at most 2.92 points), but the authors note that bound does not hold on this subset, where AgentDyn utility falls from 66.7% to 33.3%. A separate small adaptive attack search on 30 out-of-authority AgentDojo targets succeeded 21 times against the undefended agent and 0 times against PACE.

Goes deeper in: AI Agents → Security & the Lethal Trifecta → Draw the Data-Flow Graph

Related explainers

Frequently Asked Questions

Check what you knowMap your AI & GPU knowledge across every track — free, role-based