LLM·

vLLM 0.31 — SWA bounded replay — What does it mean?

The news. On October 5, 2026, the vLLM project released v0.31.0, with 717 commits from 307 contributors. For DeepSeek-V4.1-Flash, PR #56227 adds SWA bounded replay: the model's per-layer 128-token sliding-window KV cache leaves prefix caching, and after a hit the scheduler recomputes the hit's last window to rebuild it. It is on by default when vLLM uses its newer model runner (V2) without prefill context parallelism (otherwise the window stays in the prefix cache), and it is disabled on AMD ROCm GPUs (#57906); --no-swa-bounded-replay restores the old behaviour. Read the release →

Picture resuming a series from a bookmark. You do not need a recording of what you were thinking at every minute of the season: the plot notes on file carry the long story, and the only thing you are missing is the feel of the last few minutes. So you rewind a little before the bookmark, rewatch, and carry on. Bounded replay makes the same bet: keep the long-range state, and rebuild the short-range state, because it is cheaper to rebuild than to archive.

DeepSeek-V4.1 keeps several caches in each layer. Two of them, the compressed MLA cache and the sparse-attention indexer cache, cover the whole prompt, and they behave like the KV cache you already know: what was cached for the first 2,000 tokens is still valid when token 6,000 arrives, so vLLM can hash the prompt in blocks and hand a matching prefix to a new request. Those are the plot notes. The third is a 128-token sliding-window cache. Sliding-window attention is a causal mask with a lower bound as well as an upper one: each token looks back over only the last 128 tokens, so at any moment the layer needs just 128 keys and values. That is the memory of the last few minutes.

Storing that memory is where the trade goes wrong. A prefix hit can end at any block boundary, so for the window to be ready at whatever boundary H a future request stops on, the cache would have to keep sliding-window entries for every block of the prompt: a full-length record for a layer whose whole design is to need only the last 128 tokens. On any one hit, exactly one window of that record is read. The PR states the conclusion directly: storing the window for prefix caching and KV connectors costs more than it saves. It gives the conclusion rather than a measurement, and it does not say how the old path kept the window; the block-boundary argument above is our reading of why, not a figure from the PR.

Replay changes what a hit means. When the KV cache manager finds the longest matching prefix and it ends at token H, the scheduler records replay_start = H - 128 and rewinds the request's computed-token count by one window, so the 128 tokens before the hit go through prefill again. Their MLA and indexer entries are already in the cache, so the worker sends their writes for those groups to padding slots and the replayed tokens never re-enter the compressed caches; only the sliding-window cache is written. From token H onward, every token, including the first one decoded, sees a full 128-token window.

The rewatch has one honest flaw, and the metaphor already shows it: at the first rewound minute you do not remember what came just before it. The replayed tokens' window attention reads nothing below H − 128, so a replayed token near the start of the window sees a shorter window than it saw in the cold run. The PR says outright that outputs after a hit are therefore not bit-identical to a cold run, by design. Its check is GSM8K on DeepSeek-V4.1-Flash on four GB200s, run twice so the second pass hits the cache on every prompt: exact match was 0.9272 cold and 0.9212 all-hits with replay, against 0.9295 and 0.9295 without, a gap the author puts within run-to-run noise (one standard error is 0.0075 at n = 1,319). Two smaller rules follow from the design: a hit of 128 tokens or fewer is not adopted at all, since replay would recompute all of it anyway, and with replay on, KV connectors transfer only the MLA and indexer caches.

Cache group (DeepSeek-V4.1)In the prefix cache?On a hit at token HSent by KV connectors?
Compressed MLA cacheYesreused as storedYes
Sparse-attention indexer cacheYesreused as storedYes
128-token sliding-window cache, replay offYesreused as storedYes
128-token sliding-window cache, replay on (default)Norebuilt by recomputing H − 128 to HNo

Put numbers on it, holding three things fixed: a 6,144-token shared prefix that every request starts with, a 256-token new tail per request, and the 128-token window (illustrative prefix and tail sizes; the PR reports no throughput or memory figures). A cold request prefills all 6,400 tokens. A cache hit with replay off prefills only the 256-token tail. A hit with replay on prefills the tail plus one window: 384 tokens instead of 256, which still skips 94% of the cold prefill. In exchange, the sliding-window layers store nothing for that prefix. If a window were instead kept ready at every block boundary, that would mean holding up to 6,144 token entries per layer, of which any single hit reads 128: up to 48× more than a hit ever uses (an upper bound under that assumption; the PR reports no memory figures). The trade is a fixed 128 extra prefill tokens per hit, against storage that grows with every cached prefix.

Goes deeper in: Inside vLLM → KV Cache Manager → The Longest Match

Related explainers

Frequently Asked Questions

Check what you knowMap your AI & GPU knowledge across every track — free, role-based