LLM·

Allocate low-bit recurrent state precision by channel decay — Decay-aware bit allocation for recurrent states — What does it mean?

The news. On September 25, 2026, Hongren Chen and Jiayang He posted Low-Bit Recurrent States in Hybrid Language Models (arXiv 2609.30950, submitted to ICASSP 2027, not yet peer reviewed). It quantizes the recurrent state of hybrid models including Kimi-Linear, Qwen3.6-35B-A3B and Qwen3.8-27B, and reports that at a four-bit average its allocation cuts the extra prediction error by 3.3× to 27.9× compared with the best of seven baselines, with no calibration data, rotation or retraining. Read the paper →

Picture a wall of chalkboards. At every token the model fades each board a little, writes a new line on it, and then reads from it. The boards do not all fade at the same speed. Some boards keep 99% of their chalk per token and still show a mark 100 tokens later; others keep only 37% and are practically blank after one token. Now suppose you must write with chalk of limited fineness, so every write leaves a small smudge. A smudge on a fast-fading board disappears before anyone looks twice. A smudge on a slow-fading board stays, and it is read again at every one of the next hundred tokens.

That is the situation inside a hybrid model. A linear-attention or Mamba-style layer does not keep one key and value per token like the KV cache; it keeps a fixed-size state matrix that it updates in place, so its memory cost does not grow with context but is paid in full precision by default. Each row of that matrix is multiplied by a decay factor e^-s before new information is added, and in Kimi Delta Attention every row has its own s. Storing that state in low precision means rounding it at every write-back, and each rounding leaves a small error in each row. The paper's first result is that the damage from that error is set by how long the row keeps it: under its linear-noise model, for a row with constant decay s and no erasure, the error energy read back over all later tokens is proportional to how strongly queries read that row times 1 / (e^(2s) − 1).

So the paper hands out bits the way you would hand out chalk. Ordinary range-based quantization looks only at how large a row's values are, so two rows with the same range get the same bits. Decay-aware allocation gives each row a cost c = Γ̂ × R̃², where Γ̂ estimates the persistence term above from the row's decay gate and an approximate erasure rate (it tracks a diagonal entry of the observability Gramian), and R̃ is the row's value range relative to the other rows that are normalized together at the layer's output. It then applies a textbook rule: every extra bit cuts a row's squared error by 4×, so give each row the average budget plus ½·log2 of its cost divided by the typical cost (the geometric mean over all rows). The average stays fixed; slow rows get fine chalk and fast rows get thick chalk. The costs come from the model's own decay gates and state ranges at each write-back, so no calibration data, no rotation and no retraining are needed.

32-bit float — virtually continuous

01234π = 3.14159Quantized: ≈ 3.1416Δ = 0.00000Representable values (0 → 4)

The number line above is the unit of cost: fewer bits means a coarser grid and a larger rounding error for each value, as described in How Numbers Shrink. In a weight matrix that error is paid once. In a recurrent state it is paid again at every token the row still remembers, which is why the same grid can be harmless in one row and ruinous in another.

The decay rates need their own treatment. The paper also quantizes the decay gates, and it argues that an error in s matters relative to s itself: a row with s = 0.001 remembers for about 1,000 tokens, and a small absolute error there changes its memory a lot. So the gate levels are spaced evenly in log s, not in s. With that spacing, two-bit gates raised WikiText perplexity by at most 0.27 on the five models tested; with levels spaced evenly in the gate value, Kimi reached a perplexity of 86.9 even at four bits, against 9.33 in full precision. The main experiments then keep the gates at full precision and quantize only the state.

Worked example (illustrative). Use the paper's constant-decay, no-erasure proxy with equal query strength, and hold three things fixed: two rows with the same value range, an average budget of 4 bits per value, and one write-back per token. Row A has s = 0.01, so it keeps 0.99 of itself per token; row B has s = 1, so it keeps 0.37. Their persistence weights are 1 / (e^0.02 − 1) ≈ 49.5 and 1 / (e^2 − 1) ≈ 0.157, so one rounding error in row A costs about 315× as much as the same error in row B. The rule gives row A ½·log2(315) ≈ 4.15 more bits than row B, which with a 4-bit average is about 6 bits for A and 2 bits for B. Uniform 4-and-4 gives a total cost of (49.5 + 0.157) / 2^8 ≈ 0.194; 6-and-2 gives 49.5 / 2^12 + 0.157 / 2^4 ≈ 0.022. Same storage, about 9× less error under the model's own proxy. The paper warns that this proxy is optimistic: in its simulations it overstated the gain over uniform four-bit allocation by 1.7 to 2.6 times, which is why the measured results below matter more than the formula.

Per-token state quantization, 4-bit averageBest of seven baselines (excess NLL)Decay-aware (excess NLL)ReductionSource
Kimi-Linear (expert-pruned 35B-A3B)+0.131 (TurboQuant)+0.0403.3×Table 2a
Qwen3.8-27B+0.073 (TurboQuant)+0.0154.8×Table 2a
Qwen3.6-35B-A3B+2.41 (INT, block 32)+0.08627.9×Table 2a

The Qwen3.6 row is the sharpest. There, every four-bit baseline did worse than simply wiping the state to zero every 64 tokens (+0.373 nats), while decay-aware allocation stayed at +0.086. At six bits, the paper reports excess NLL of at most 0.0034 nats, almost indistinguishable from an FP32 state. For scale, a 4-bit payload is one eighth of FP32 storage before metadata; including per-row scales and width maps, the paper counts 4.156 bits per value on Kimi and 4.5 on the Qwen models.

The limits are stated plainly. The advantage shrinks when the state is written back less often: with one write-back per 64-token chunk the method was best in only 4 of 15 model and budget pairs, tied in 7 and worse in 4, because fewer roundings leave less error to accumulate. The quantizers were simulated, with no kernel timings, so the paper makes no speed claim. And the error model is an approximation that ignores how noise travels between layers.

Goes deeper in: LLM Internals → Quantization → What to Quantize

Related explainers

Frequently Asked Questions

Check what you knowMap your AI & GPU knowledge across every track — free, role-based