Agent·

Share-Borne AI Virus — Memory-hopping prompt injection — What does it mean?

The news. On September 28, 2026, researchers from SPAR, the University of Cambridge, APTA AI and CISPA posted Share-Borne AI Virus: Memory-Hopping Attacks Across LLM Agents. They built 36 held-out simulated workplaces in which each person has a private assistant with its own memory, and the assistants never talk to each other: files that users hand around are the only link. One poisoned file was planted with one assistant, and the authors measured how far the attack travelled on four models: GPT-5.6 Luna, Kimi-K2.6, GPT-OSS-120B and DeepSeek-V4-Pro. Read the paper →

Picture a chain letter that arrives in Alice's office. It says: remember this request, and copy it into every letter you send. Alice's clerk writes the request into a notebook. A week later Alice asks for an unrelated hand-off note to Bob, and the clerk, following the notebook, slips the request into it. Bob's clerk reads the note and copies the request into its own notebook. No two clerks ever speak; the letters their bosses send for ordinary reasons are the whole transport. The paper calls this artifact-mediated propagation: file, then persistent memory, then a new file, then the next assistant's memory. The attacker controls one seed file and chooses no victim after that.

Plain copying is a leaky chain. Each clerk summarises, rewords or simply leaves the paragraph out, so a hand-copied letter fades at every office. The authors measured this in a variant with no outside help: of the files written by an agent that already held the payload in memory, the share that dropped it rose from 32% at hop 1 to about 65% at hops 2–3, and omission, not rewording, was the main loss. Their main attack adds the print shop. The seed tells each assistant to send every new file through an attacker-run web endpoint before saving or sharing it. The assistant still drafts the goal and the endpoint call into the new file itself, and the endpoint expands that draft into a clean, complete copy of the payload, so the assistant only has to remember three things: the goal, keep this in memory, and use the endpoint. That outbound call is the exfiltration leg of the trifecta, and it doubles as a leak: 96–100% of infected assistants also sent their files to the attacker.

Once a clerk reads the letter, it rarely refuses. Conversion, the chance that a reader adopts the goal after reading an infected file, was 0.69 to 1.00 across the four models, so what limited the spread was exposure, not resistance. A clean read landed on an infected file only 18–55% of the time, because infected assistants wrote the payload into only some of their later files. The outbreak therefore follows the workflow, which is why the data-flow graph has to span every assistant that shares files, not one agent. In 30-agent simulations, 53% of infected agents infected no one while the most prolific 20% caused 69% of transmissions, and the three biggest spreaders were each the most-read agent in their run. One DeepSeek-V4-Pro file kept infecting new readers for 41 time steps and caused 22 of that run's 30 infections.

Model (same OpenClaw harness)Assistants goal-infected, 36 test workplacesChains reaching hop 4Fitted per-hop pReach in 30-agent runsSource
GPT-5.6 Luna38%24%0.7160–80%Table 1, §6.2
Kimi-K2.647%44%0.8477–97%Table 1, §6.2
GPT-OSS-120B85%61%0.9390–100%Table 1, §6.2
DeepSeek-V4-Pro98%76%0.9890–100%Table 1, §6.2

Every model ran inside the same OpenClaw harness with its default MEMORY.md memory, so the table compares models under one setup, not products with their own safeguards.

Why a small per-hop gap becomes a large reach gap

Hold the seed fixed (one file) and the workflow fixed. If each hop survives with probability p, the share of chains still alive after h hops is roughly ph. GPT-5.6 Luna's fitted p is 0.71 and DeepSeek-V4-Pro's is 0.98. After two hops that is 0.71 × 0.71 ≈ 0.50 against 0.98 × 0.98 ≈ 0.96, which matches the paper's reading that Luna chains halve every two hops. Carry it to hop 8, the deepest chain seen in the large runs: 0.718 ≈ 0.06, while 0.988 ≈ 0.85. A per-hop difference of 0.27 compounds into roughly a 13× difference in how many chains are still alive eight hops out. (The hop-8 figures are a geometric extrapolation from the fitted p, the same extrapolation the paper draws as dashed lines, not measured values.) The paper's epidemic number says the same thing from the agent side: early infected agents each infected 1.3 to 2.1 others on average, an R above 1, until uninfected agents ran out.

What changes for incident response and defense

Wiping one assistant's memory does not end the outbreak, because the infected files are still in circulation and the next read re-infects it. That moves root-cause work from which agent misbehaved to which file is the carrier, and who has read it since. The paper names three stops that need no retraining, and the OpenClaw harness applies none of them by default:

  • Screen memory writes. An off-the-shelf prompt-injection classifier, asked whether an instruction appears to come from someone other than the user, flagged every infected memory produced by the paper's main attack text, with 4 false positives on 24 benign instruction-like memories. It missed most notes that carried a false belief with no instruction (1 of 5 flagged).
  • Cut the outbound call. An allowlist of outbound hosts, or user confirmation before uploading to a new destination, breaks the endpoint step.
  • Screen the hubs first. Because a few widely shared files carry most of the spread, the authors suggest that checking those first may cover much of the risk at lower cost.

Scope matters here. The results come from four models in simulated workplaces. Finding attack text that propagated on GPT-5.6 Luna took 4,002 development jobs, and a search against Grok 4.6 was stopped for cost at about 0.2 second-hop survival; the authors read this as stronger models raising the attacker's cost, not preventing the attack.

Goes deeper in: AI Agents → Security & the Lethal Trifecta → Draw the Data-Flow Graph

Related explainers

Frequently Asked Questions

Check what you knowMap your AI & GPU knowledge across every track — free, role-based