KV-streams — Keeping the KV cache through context compaction — What does it mean?
The news. On September 28, 2026, a group of 18 researchers posted KV-streams for Efficient Compaction in Agentic Reinforcement Learning on arXiv. It is a plug-in for the inference engine and the trainer that is designed to work with any compaction strategy that fits the paper's delete-and-optionally-summarize definition. The authors report a 2.6× to 5× wall-clock speedup in training across the compaction strategies they tested, no evidence of lower task scores, and a finding that the kept cache can carry information from deleted text after RL alone, with no supervised fine-tuning stage first. Read the paper →
Why a fresh board costs you twice
Picture a whiteboard that fills up during a long meeting. The usual fix is to take a clean board, copy over the lines you still need, and throw the old board away. The copying is the expensive part: every line you keep is written a second time, and you do it again at every clean-up.
An agent's KV cache is that whiteboard. It grows with every token, so one long rollout can take GPU memory that several short ones could use (why it grows). Compaction caps it: when the context reaches a budget, some turns are deleted and maybe replaced with a summary. Most pipelines then start a fresh trace — a new, separate sequence — with the kept text. The kept tokens now sit at new positions, and because a cached key and value depend on position, their old cache entries no longer match. The model must prefill them again — in the inference engine that samples the rollout, and again in the trainer that runs the forward and backward pass.
Sponging off only the old section
KV-streams does what you would do with a sponge: wipe the section you no longer need and keep writing on the same board. In the inference engine it deletes the old span's cache entries in place, so the kept entries keep their original positions and are never recomputed. Generation then continues on one unbroken stream of cache entries.
The trainer has to see the same thing the sampler saw, or the gradient would be computed on a different sequence. Instead of splitting the rollout into several traces, the trainer runs one continuous trace with an attention mask that blocks the deleted span, which reproduces the deletion in a single forward pass. The paper applies this to four compaction strategies, from keeping almost nothing to keeping almost the whole budget.
Worked example (illustrative). The paper gives the extra prefill of re-prefill compaction as N·r/(B−r), where N is the rollout length, B the token budget and r the tokens kept at each compaction. Hold N = 200,000 tokens and B = 32,000 tokens fixed (32K is the paper's TextWorld budget; the rollout length is illustrative) and change only what is kept. Keep an 8,000-token summary: room for 24,000 new tokens between compactions, about 8 compactions, and about 67,000 extra tokens (N/3, the paper's own B/4 case). Keep half, 16,000 tokens: 12.5 compactions and 200,000 extra tokens, a second full rollout. Keep a 28,000-token sliding window: a compaction every 4,000 tokens, 50 of them, and 1,400,000 extra tokens — 7× the rollout itself. With KV-streams, the repeated prefill of kept tokens is 0 in every case; implementation padding can still add tokens. These are token counts, not time: generation and other work do not shrink, which is one reason the measured speedups are smaller.
| Strategy | What each compaction keeps | Re-prefill cost (paper's analysis) |
|---|---|---|
| Summary | Only a model-written summary | Low; N/3 when the summary is B/4 (source) |
| Markovian Thinker | The most recent half of the context | Medium; r = B/2 gives N extra (source) |
| Markovian Pick | Turns the model chooses to keep | Depends on how much the model keeps |
| Sliding window | Almost the whole budget | Highest; grows without bound as r nears B (source) |
What the measurements show
On TextWorld, KV-streams reached the final score of re-prefill compaction with 3.7× to 11.3× fewer GPU-hours (paper). On ALFWorld, where traces are shorter (a 16K budget), the throughput gain was smaller, from 1.3× to 3.0× across the four strategies. On software-engineering tasks with Qwen3.5-4B, sliding window with KV-streams scored 52.7 ± 2.1% on SWE-bench Verified, within the margin of error of full context (51.5 ± 1.2%) and re-prefill Markovian Thinker (53.4 ± 0.2%). The speed difference was large: about 20 hours to peak versus more than 65 hours for full context. The authors note that TextWorld and the software runs used a single seed, and only ALFWorld was run with several seeds.
The engineering has a catch. vLLM caches the KV in 16-token blocks and adds only complete blocks to its prefix cache, so a completion that ended mid-block broke the stream. The authors pad every completion to a multiple of 16 tokens, which costs extra tokens; SGLang, with a block size of 1, needs no padding.
The smudge that remembers
Back to the whiteboard. Lines you wrote while the old section was still visible were written with that section in view, so they may still hold its gist after you erase it. The kept cache entries work the same way: they were computed while the deleted text was still in view, so they can carry its information forward. Earlier work saw this and needed a supervised fine-tuning stage to make it appear. In a controlled test — assign an object, delete that turn, then ask which object it was — RL alone reached 100% recall once the kept cache was large enough, and recall was unreliable only at the smallest budgets of 16 and 32 tokens.
This is also the tradeoff to keep in mind. With KV-streams, what the model knows is no longer the same as the text in its window: a fresh prefill of the kept text would give different cache entries from the ones KV-streams keeps. That can help recall, but it also means you cannot read the model's visible context and know everything it can use. It is a different choice from text-level rules like CliffCompaction, which decide which words survive; KV-streams decides whether the cache behind those words survives.
Goes deeper in: LLM Internals → KV Cache → Prefill vs Decode