CacheReforge — Stale KV cache repair after LoRA adapter updates — What does it mean?
The news. On September 25, 2026, Yuhang Cao, Yanzhou Mu, Chunrong Fang and Zhenyu Chen posted CacheReforge: Bounded Recovery for Stale KV Caches under Evolving Adapters to arXiv. They freeze Qwen2.5-1.5B and Qwen2.5-7B, keep updating a rank-8 LoRA adapter, and ask how much of a long prompt's cached K/V must be recomputed after each update. On 16K-token HotpotQA and 2WikiMQA with 32 consecutive updates, the method cuts mean KL by 92.4% against stale reuse while recomputing 5.44% of layers. Read the paper →
Picture a building drawn as a stack of floor plans, one sheet per floor, where each floor's drawing is worked out from the one below it. The architect revises floor 10. The sheets for floors 1 to 9 are still correct. Every sheet above 10 might now be wrong, because each one was derived from a floor that changed. A cached prompt is exactly this stack: one sheet of key and value vectors per layer, each computed from the output of the layer below.
A normal KV cache never faces this problem, because the weights never change between writing a sheet and reading it. LoRA breaks that assumption when the adapter keeps training: the tokens in the prompt are identical, but the projection matrices that turned those tokens into K and V are not. A prefix cache keyed on token blocks cannot see the difference. The paper notes that existing systems judge validity from token and context identity or a fixed adapter identity, so a cache from version 3 of an adapter and one from version 4 look the same.
Which sheet goes stale first depends on what the update touched, and this is the most concrete mechanism in the paper. If the update changes the key or value projection of layer 10, layer 10's own cached K/V are already wrong. If it changes the query projection or the MLP of layer 10, layer 10's K/V are still correct — they never used those weights — and the first wrong sheet is layer 11, which reads layer 10's changed output. The paper confirms this directly: for query and MLP updates, recomputing only the changed layer repairs nothing and matches stale reuse on every test row.
You cannot patch a sheet in place by applying the weight change to the stored vectors. Every sheet above the revision was computed from a hidden state that has also changed, so a correct patch would mean running those layers again anyway. Recovery therefore means recomputing a contiguous range of layers, starting from a saved copy of the hidden state just below the change — the drawing of floor 9 you kept for exactly this purpose.
The obvious safe rule is to reprint every floor from the revision to the roof. The paper's central finding is that this is often more than needed and sometimes not enough to predict. How far an update can spread (the dependency depth) is a different number from how far you must recompute before the output is correct again (the recomputation horizon). Across 72 test profiles on Qwen2.5-1.5B, stale reuse was already safe in 37, a short bounded window was enough in 18, and 17 needed the whole affected suffix. The error also does not fade smoothly: in 50 of the 72 profiles the error went up at some point as the window grew, and hidden-state drift often stayed small through the middle of the network and peaked near the final layers.
So CacheReforge keeps a revision stamp on every sheet. Each cached layer keeps a copy of the adapter state that produced it — its anchor, a set of weights rather than just a version number — so a cache that has been partly repaired several times is a mixed-version object, not a single version. After each update it measures how far the current adapter has moved from each layer's anchor, multiplies that by a per-layer sensitivity calibrated offline, and recomputes up to the furthest layer whose risk crosses a threshold. Layers left untouched keep their old stamps, so drift that builds up over several small updates is still caught later. The paper also shows why it looks at the whole unrepaired tail rather than only the layer where recomputation stops: the summed error of the remaining stale layers tracked output KL with a Spearman rank correlation of 0.873 (1.0 would be a perfect ranking), against 0.666 for the error at the stopping layer alone.
| Policy after each update | Layers recomputed | Mean KL vs fresh | Cache upkeep (s) | Source |
|---|---|---|---|---|
| Fresh full prefill | 100% | 0 | 50.26 | Table 1 |
| Exact affected suffix | 50% | 0 | 28.75 | Table 1 |
| CacheReforge | 5.44% | 0.0335 | 3.42 | Table 1 |
| Stale reuse | 0% | 0.4385 | 0.31 | Table 1 |
Walk the numbers in the paper's main test: Qwen2.5-7B, 28 layers, 16K-token prompts, and 32 consecutive updates, each to the value projection of a middle layer. For every one of those updates the exact safe answer recomputes the affected suffix, which is 14 of 28 layers (50%). Averaged over the held-out run, CacheReforge recomputed 5.44% of layers, which is about 1.5 of 28 layers per update (5.44% × 28, our arithmetic) — an average across updates, not a fixed window. Over the 32-update run that took cache upkeep from 50.26 s for fresh prefill to 3.42 s, a 93.2% cut, and mean request time per update from 4.16 s to 2.30 s. The price is that the result is close to fresh, not equal to it: mean KL is 0.0335 instead of 0, and top-1 agreement (how often the most likely next token matches the fresh-prefill one) is 94.04%, up from 84.47% for stale reuse. In short, about one and a half recomputed layers per update removed 92.4% of the output drift that stale reuse caused, at under a tenth of the upkeep cost of a full prefill. Task scores moved the same way but less completely: HotpotQA token-F1 went from 7.95% (stale) to 9.13%, against 9.41% for fresh prefill.
Two limits keep this a research result rather than a serving feature. The evidence comes from two Qwen2.5 models and a single rank-8 adapter, on 4K and 16K controlled workloads plus 16K question-answering tasks, and the end-to-end cost result covers only middle-layer value-projection updates; the sensitivity profile is calibrated per deployment configuration, and the method stores extra hidden states at restart points that a plain KV cache does not keep. The paper cites prefix-reuse systems such as Prompt Cache and SGLang as motivation, not as tested baselines, and argues that systems of this kind cannot represent a cache written by an earlier version of the same adapter. The durable lesson is independent of the method: a cached K/V entry is valid for a (tokens, weights) pair, not for tokens alone, so any system that changes weights under a live cache — adapter updates, online training, hot-swapped fine-tunes — needs a version in the cache key and a plan for what to recompute. That is also why prefix caching wins or loses on how often its assumptions hold.
Goes deeper in: LLM Serving → Prefix Caching & RadixAttention → Block-Hash Chain (vLLM APC)
Related explainers
- Selective KV recomputation for RAG chunks — the same choice of what to recompute, but when the context around a cached chunk changes instead of the weights
- Prefill-only LoRA adapters — a design where the adapter's effect lives entirely in the prefilled KV cache, which is exactly the state that goes stale here