LLM·

ResidualQuant stores looped-Transformer KV caches as 2-bit residuals — Cross-loop KV residual quantization — What does it mean?

The news. On October 7, 2026, researchers from KAIST, Yonsei University and Seoul National University posted "ResidualQuant: KV Cache Quantization for Looped Transformers with 2-Bit Residuals" on arXiv. They test it on two looped models, Ouro-1.4B and Huginn-3.5B, on math reasoning and code generation benchmarks, and integrate it into vLLM with a custom CUDA attention kernel. On one RTX 5090 they report up to 2.73× higher decode throughput at a fixed batch and up to 4.15× higher peak throughput. Read the paper →

Picture four photos of the same mountain, taken a few seconds apart. Storing four full prints wastes space, because the prints are almost identical. A careful photo lab keeps one sharp master print and, for each other photo, a thin tracing-paper overlay that marks only what changed. The overlay holds very little detail, so a cheap pen is enough to draw it.

A looped Transformer creates exactly this situation. It runs the same shared block several times for each token, and each run (a loop) computes keys and values with the same projection weights. When the hidden state changes only a little from one loop to the next, the keys and values change only a little too. Yet the cache still stores a full copy per loop, so KV cache memory grows with loop count on top of context length and batch size. The paper shows that the residual between loops has far fewer large outliers than the raw keys. Outliers are what make low-bit quantization fail: one big value stretches the scale for its whole group and leaves the small values with almost no levels.

So ResidualQuant keeps the final loop's KV as an INT4 master print and stores every earlier loop as a 2-bit overlay: residual = KV(loop) − c × anchor, where c is a least-squares coefficient computed for each key or value vector. Two refinements shrink the overlay before it is quantized. The coefficient c corrects for KV magnitudes that differ across loops, and a learned rotation spreads the remaining outliers across channels. At attention time a custom CUDA kernel rebuilds each loop's KV as c × anchor + residual, tile by tile in on-chip memory, so the rebuilt tensor is never written back to HBM.

Outliers force INT8 range to span –60 to 60 · most of the 256 grid slots fall on empty space

Weight value distribution-60-40-200204060
99.9%
0.1%
outliers
INT8 grid:-60 to 60 — 256 slots wasted on empty space
■ main distribution■ outliers■ INT8 grid

Which print is the master matters for speed as much as for accuracy. The obvious design chains the overlays: loop 2 relative to loop 1, loop 3 relative to loop 2, and so on. Then rebuilding loop 4 of a 4-loop model means reading the loop-1 anchor plus three residuals, which is four KV reads where plain quantization needs one. Decode spends most of its time reading the cache, so those extra reads cost throughput directly; the paper measures shared-anchor reconstruction at up to 1.62× faster than the chained version. With one shared master, each other loop needs two reads: the anchor and its own residual. The authors pick the last loop rather than the first as that master, because it gave lower reconstruction error in the early loops on Ouro-1.4B. The cost is timing: the current token's master does not exist until its final loop finishes, so that token's KV stays in BF16 in a temporary buffer and is quantized only after the last loop.

Each refinement adds accuracy on its own. The ladder below is the paper's single-model run (Ouro-1.4B on the MATH500 math benchmark, INT2 groups of 32 channels), so read the numbers as one configuration, not a general law.

ConfigurationMATH500 accuracyAvg. KV bits (excl. metadata)Source
Direct INT2, every loop27.8%2paper, Fig. 1
plus last-loop residual57.2%2paper, Fig. 1
plus least-squares scaling60.6%2paper, Fig. 1
plus rotation of the residual71.6%2paper, Fig. 1
plus INT4 anchor (full ResidualQuant)76.0%2.5paper, Fig. 1
BF16 baseline75.0%16paper, Fig. 1

Worked example: where the 80.7% comes from. Hold the model fixed at Ouro-1.4B with its four loops and count bits for one KV element position. Uncompressed, each loop stores BF16, so the four loops cost 4 × 16 = 64 bits. ResidualQuant stores the anchor at 4 bits and the three residuals at 2 bits each: 4 + 2 + 2 + 2 = 10 bits, an average of 2.5 bits per element, or 6.4× smaller before metadata. Each group of channels also stores an FP8 scale and offset, and each residual key or value vector stores one BF16 coefficient; with that metadata the paper's ratio is about 5.2× smaller, its reported 80.7% cut. Now hold the GPU (one RTX 5090) and context (16k tokens) fixed. At batch 2, decode goes from 34.0 to 93.1 tokens/s (2.73×), because each step reads less KV from memory. The freed memory also raises the largest tested power-of-two batch that fits from 2 to 4, which reaches 141.3 tokens/s, 4.15× the BF16 peak.

The method is built for looped Transformers, where shared projection weights make the loops' KV similar; the paper does not test ordinary Transformers. In a standard model every layer has its own weights, so there is no obvious near-duplicate to subtract, and whether a similar trick helps there is an open question. Within its scope, the evaluation covers two looped models on math and code benchmarks; across those model–benchmark pairs the mixed-precision configuration averages 52.37 against BF16's 52.94, and the throughput results come from one consumer GPU. Earlier designs such as MoR, PLT and MELT share one loop's KV across loops, which saves memory but can discard loop-specific information; ResidualQuant keeps a distinct KV for every loop, only at lower precision.

Goes deeper in: LLM Internals → KV Cache → Memory Cost

Related explainers

Frequently Asked Questions

Check what you knowMap your AI & GPU knowledge across every track — free, role-based