LLM·

Pretrained LLMs resist 4-bit quantization partly because layer errors cancel — Counteracting quantization error — What does it mean?

The news. On September 10, 2026, Yuxiang Chen, Michael Beyer, Jun Zhu and Jianfei Chen posted Why Does Post-Training Quantization Work? — an analysis paper rather than a new quantizer. It starts from an embarrassment. Casting Qwen3-32B to the 4-bit NVFP4 format by plain round-to-nearest, with no calibration at all, costs only 0.43 percentage points of average accuracy across six zero-shot benchmarks, even though the model never saw quantization noise during training. Rounding every linear weight in every block ought to corrupt the hidden state more and more with depth. Instead the discrepancy grows far more slowly than that prediction suggests, and the paper sets out to measure why. Read the paper →

Picture the line of scribes again. Every one of them misreads a letter here and there, and the page they hand on carries those slips forward. If the slips were independent, the last scribe would be holding something close to nonsense: each new error lands on top of every error already there, and the damage grows with the length of the line. That is the naive prediction for a quantized transformer, and it is roughly what happens when the scribes cannot read the language.

The paper's first finding is that in a pretrained model the slips are not independent — the error a block newly introduces tends to point against the error it inherited from its input. A fluent scribe who receives a slightly wrong sentence tends not to carry the wrongness forward faithfully — though the analogy stops at the outcome, because no block notices an error or repairs it, and what the paper establishes is a statistical tendency rather than an act of correction. The two errors partially cancel, so the gap between the full-precision pass and the quantized pass widens across the residual stream far more slowly than a per-layer error budget would predict.

The crucial detail is what distinguishes the two cases. It is not how faithfully each individual weight survived rounding, because that is matched. The authors show the difference is something the model acquires during pretraining: measured across OLMo3 and Pythia checkpoints, counteraction is weak near random initialization and becomes more pronounced at later checkpoints. Nobody put it there deliberately; it arrives as a side effect of ordinary pretraining, which is what makes it surprising.

32-bit float — virtually continuous

01234π = 3.14159Quantized: ≈ 3.1416Δ = 0.00000Representable values (0 → 4)

Every one of those slips starts as the ordinary rounding above: a weight lands between two representable values and gets snapped to the nearer one. The question is what happens after that error is created, and the cleanest way to see it is to hold almost everything fixed and change exactly one thing, which is what the paper does.

Fix the model shape (Qwen3-32B). Fix the format and the rounding rule (NVFP4, round-to-nearest, no calibration). Fix how faithfully each individual weight is stored — quantized weights sit at a cosine similarity of about 0.996 to their full-precision values, and the paper reports that a randomly initialized model's weights round to nearly the same fidelity. Now change the one remaining variable: run the forward pass from the pretrained checkpoint, then from random initialization.

If weight-level closeness were the explanation, both runs would land in the same place. They do not. The randomly initialized model finishes with 5.5× the absolute hidden-state error and 6.7× the relative hidden-state error of the pretrained checkpoint — same per-weight fidelity, same rounding rule, same depth, several times the damage. So per-weight reconstruction fidelity cannot be what protects the pretrained model. What pretraining changes is how the resulting errors interact as they propagate — whether they line up with one another or against. The same pattern holds for OLMo3-7B against a step-0 checkpoint, at 20.2× the final absolute error and 3.7× the relative error. And when the authors intervene to remove counteraction directly, fast error growth returns — causal evidence rather than a correlation.

Slowing the error down is only half the answer, because some error does survive to the end of the stack. The second half is about what kind of error survives, and who reads it.

Split the surviving discrepancy into a length part and an angle part, and almost all of it is angular — 88.9% to 98.7% of the squared relative hidden error across depth — while the hidden state's norm barely moves. The final hidden state is not stretched or shrunk so much as turned. After the final normalization that turn averages 12.47°.

Then it reaches the LM head, a very tall matrix acting in a very high-dimensional space, and the geometry there is forgiving in a specific way. In high dimensions the rotation is strongly attenuated when measured as the change in angle to a fixed LM-head row, and empirically it disturbs least the rows the model was already most confident about. Higher-ranked tokens have smaller projection angles to the hidden state and correspondingly smaller relative score errors, so the ordering near the top survives even though the vector moved.

The output numbers follow from that. Across three datasets the top-ranked token changes 8.3% to 12.7% of the time, while roughly 85% of the top-10 and top-20 sets are retained on average. The predictions the model is most confident about are the ones quantization is least able to disturb, which is how a substantial rotation of the final hidden state still produces only small changes at the output — and why a benchmark average can move so little.

Proposed explanationDoes it survive the paper's test?Evidence
Quantized weights stay close to the originalsNot on its ownA randomly initialized model rounds to nearly the same per-weight fidelity (cosine ~0.996) and still accumulates several times the hidden-state error — §1
Counteracting residual errorYes — identified as a major factorAn exact decomposition of error growth, plus interventions that remove counteraction and restore fast growth — §3
LM-head geometry absorbs what is leftYes, at the output stageThe surviving error is almost entirely a rotation, and top-ranked token scores move least under it — §4

Three cautions before this reads as a licence. The experiments measure the numerical effect of quantization, not the runtime or memory saving of a packed 4-bit kernel — this is a paper about why the outputs survive, not about how much faster anything runs. The paper's setup quantizes the linear projections inside attention and the MLP and leaves embeddings, normalization layers, activations and the LM head in BF16 — the forgiving geometry the second mechanism depends on is full-precision geometry, so a recipe that also quantizes the head is not covered by this result. And the finding explains why a well-pretrained model tolerates rounding in the formats and settings the authors tested; it says little about a checkpoint with far less pretraining behind it, or about a recipe outside that tested range. Quantization has other failure modes — the outlier problem above all — that this paper does not speak to. Counteraction is something pretraining builds up, so it is reasonable to expect less of it wherever there has been less pretraining.

What changes in practice is mostly how you reason about the risk. The comfortable rule of thumb — keep each weight close enough to its original and the rest takes care of itself — turns out to be measuring the wrong thing, since the failing case has that property too. The quantity that actually tracks output damage is how the error grows across depth, and after that, how much of it the LM head's geometry can absorb.

Goes deeper in: LLM Internals → Quantization → The Quantization Process

Related explainers

Frequently Asked Questions

Check what you knowMap your AI & GPU knowledge across every track — free, role-based