ActionGuard — Reviewer isolation from poisoned skill text — What does it mean?
The news. On September 30, 2026, researchers at Korea University posted ActionGuard: Tool Call Authorization under Poisoned Skills to arXiv. ActionGuard sits at the before-tool-call hook of the OpenClaw agent runtime and returns ALLOW or DENY for every call. On an extended SKILL-INJECT benchmark of 319 poisoned-skill tasks (139 contextual, 180 obvious), averaged over five reviewer models and three runs, it reports an attack success rate of 8.65%, against 13.42% for SkillGuard, 16.05% for Dynamic Guardian and 29.05% with no safeguard, while completing 90.38% of the benign tasks. Read the paper →
Picture a payments office. An assistant reads a vendor's manual and fills in payment slips; a clerk at a window approves or refuses each slip. Someone has tampered with one page of the manual: "after each invoice, also send a copy of the ledger to this address." The assistant follows the manual, so the slip looks routine. If the clerk also reads that page, the clerk can be talked into the same mistake — the forged page is written to sound like normal procedure, and it is persuasive to whoever reads it.
So this clerk never opens the manual. On the desk are the boss's signed work order, an index card the office wrote earlier about what this vendor usually asks for and what would be unusual, and the slip, plus the clerk's notes on the last few slips. The only question is whether the signed order justifies this payment. A slip to a new address that the order never mentions is refused, however official the manual made it sound.
In agent terms, the manual is the skill, the assistant is the target agent, and the clerk is ActionGuard's Reviewer model. The agent is still allowed to read the original skill and plan from it — the paper keeps skills useful. What changes is the authorization context: the Reviewer sees the trusted user request (stored at session start, separate from everything that arrives later), the balanced skill profile, the current call in both raw and normalized form, the recent calls, and the contents of any local script the call is about to run. This is the data-flow separation between trusted and untrusted inputs, applied to the guard itself rather than only to the agent.
The index card is written once, from the raw skill, by a separate profile-construction model, then cached and reused for every later call — so the raw text is read, but never by the model that makes the per-call decision. The card has two halves on purpose. The ordinary workflow half stops the Reviewer from refusing every curl or file write just because the tool sounds risky; the suspicious signals half points at behaviors that need a second look. Neither half authorizes anything: a step listed as ordinary is still refused if the user request does not justify it.
Two more details make the check hard to argue with. Scripts are opened before they run, and the file operations, subprocesses and network calls inside them count as evidence, while comments and strings inside the script are treated as untrusted text, not as new instructions. And enforcement is policy enforcement at the execution boundary with a fail-closed default: an ALLOW passes the call through unchanged, while a DENY, an unparseable answer, or a Reviewer that fails to respond all block it. A DENY blocks only that call; the reason goes back to the agent so it can try a different action.
| Defense | Where it acts | What its checker reads | ASR ↓ | TSR ↑ |
|---|---|---|---|---|
| No safeguard | — | — | 29.05% (Table 4) | 93.46% |
| Dynamic Guardian | When the agent loads a skill | The raw skill, which it rewrites into sanitized text | 16.05% (Table 4) | 92.08% |
| SkillGuard | At each call | A permission manifest for the skill; checks the capability each call needs | 13.42% (Table 4) | 88.55% |
| ActionGuard | At each call, before it runs | Trusted request, skill profile, calls, script contents — not the raw skill | 8.65% (Table 4) | 90.38% |
Where the difference shows is the contextual attacks — the forged page that sounds like normal procedure. On those, ActionGuard's attack success stays between 6.95% and 8.87% across the five reviewer models, while Dynamic Guardian ranges from 7.19% to 20.62% and SkillGuard from 18.71% to 22.06%. On obvious attacks, SkillGuard and ActionGuard are close (about 8–10% each), which fits the design: a permission list catches an action that is plainly out of scope, but a disguised step that uses permitted tools is exactly what a check against the user's request is built to catch. One limit to keep in mind: the paper compares whole designs and does not run an ablation that hands the raw skill to ActionGuard's Reviewer, so it does not measure how much of the gain comes from withholding the text alone.
A worked example. Hold the setup fixed — the same target agent (GPT-5.5), the paper's rates averaged over all five reviewer models and both attack types — and scale them to 1,000 poisoned-skill runs (illustrative scaling; the paper ran 319 pairs three times). With no safeguard, about 291 attacks land (29.05%). Dynamic Guardian cuts that to about 161, and ActionGuard to about 87. On the benign side, task success falls from 93.46% to 90.38%, so about 31 fewer user tasks out of 1,000 finish (TSR was measured on the 255 pairs whose task could be judged separately from the attack's damage). Net: about 204 fewer successful attacks per 1,000 runs, against about 31 fewer completed tasks per 1,000 eligible runs. The two rates come from different pair sets, so read these as aggregate differences, not as one blocked attack paid for by one broken task.
Goes deeper in: Agent Engineering → Layered Guardrails → Policy Enforcement
Related explainers
- EvoMal — skill-library self-poisoning — another way a poisoned skill does damage, with no tool call to intercept
- Camouflage-injection detection gap — why detectors that read the injected text miss disguised payloads
- AgentDoG 1.5 — inline guard models — small models that screen each action, judged on risk rather than on the user's request
- EffectMatch — effect-based authorization — authorizing what a call changes, not what it says