ZIP-SR 4-bit AdamW — Preconditioner-space stochastic rounding — What does it mean?
The news. On October 8, 2026, researchers from UC Berkeley and Nubank posted Rounding in Preconditioner Space: Redesigning 4-bit AdamW Optimizer-State Quantization (arXiv 2610.12444). It proposes two 4-bit AdamW recipes, ZIP-SR and ZE-EDEN, that store both optimizer moments at the same cost as TorchAO's 4-bit AdamW. Across GPT-style and Llama-style pretraining from 130M to 2.7B parameters, both shrink TorchAO's validation-loss gap to 32-bit AdamW at every size, by up to 70%. PyTorch code is on GitHub. Read the paper →
Picture a car whose cruise control sets its speed from a road-roughness gauge: rough road, slow down; smooth road, speed up. AdamW works the same way. For every weight it keeps a running average of the squared gradient, the second moment v, and divides that weight's step by √v + ε. Holding the first moment fixed, a weight whose gradients have been large gets a small step multiplier, and a weight whose gradients have been tiny gets a big one. The multiplier 1/(√v + ε) (written here without AdamW's early-training bias correction, which fades to about 1) is called the preconditioner — it is the speed setting, and it is the only way v ever touches a weight.
32-bit AdamW keeps two FP32 numbers per weight, the first moment m and the second moment v, which is 8 bytes per parameter before gradients or activations are counted (the memory side of mixed-precision training). Storing both in 4 bits, with one FP32 scale per 128-value block, cuts that to 1.0625 bytes per parameter on the tensors it covers (normalization-layer states stay in FP32) — a 7.53× reduction, about 19.2 GB saved on the paper's 2.7B-parameter pretraining model. The price is a coarse gauge: every stored v is snapped to one of 16 levels. The dangerous level is the bottom one: if a small v is rounded down to zero and the next gradient is also tiny, little more than ε is left in the divisor, and that weight's next step multiplier can be about 100× too large. Earlier 4-bit work called this the zero-point failure. TorchAO's 4-bit AdamW avoids it by removing zero from the codebook, so the bottom notch is a small positive number — but that floor makes tiny v values look larger than they are, which caps the big steps those weights should be taking.
32-bit float — virtually continuous
Back to the car. The usual way to round randomly is stochastic rounding: if the true reading sits a quarter of the way up from notch a to notch b, store b with 25% probability and a with 75%, so the average stored reading is exact. That is fair in roughness units. But the driver never reads roughness — the driver reads speed, and speed is one over a square root. A reading near the zero notch is likely to be stored as zero, and zero means top speed. ZIP-SR keeps the zero notch but sets the rounding odds in speed units: it picks the probabilities so the average multiplier 1/(√v + ε) is exact at storage time, rather than the average v. Because the zero notch sits at a huge multiplier (1/ε), matching the average multiplier makes zero very unlikely. In the paper's near-zero analysis, this keeps the next step's average multiplier error bounded, without deleting zero from the codebook — though after the next moment update the average is not guaranteed to be exactly right. The name spells this out: Zero-Inclusive, Preconditioner-space Stochastic Rounding. In a controlled 834M-parameter GPT-style run with FP32 first moments and the same zero-inclusive codebook, both state-space rules (round-to-nearest and stochastic rounding) became unstable, while both preconditioner-space rules closely tracked 32-bit AdamW.
Here is one weight whose gradients have gone quiet, with illustrative numbers: ε = 10-8 and β2 = 0.95 (how much of the old v each step keeps) are the paper's settings, the rest are chosen for easy arithmetic, and AdamW's early-training bias correction is ignored because it is close to 1 late in training. The true v is 10-12, and the two nearest 4-bit levels are 0 and b = 4 × 10-12. If the next gradient is about zero, the next-step multiplier is about 1.0 million when v is stored exactly, about 0.51 million when b is stored, and exactly 1/ε = 100 million when zero is stored — the 100× jump. State-space stochastic rounding stores zero with probability (b − v)/b = 75%, so the expected multiplier is about 0.75 × 100M + 0.25 × 0.51M ≈ 75 million. Rounding fairly in v leaves this weight's step about 74× too large on average. Preconditioner-space stochastic rounding stores zero with probability ε(1 − √(v/b)) / (√v + ε) ≈ 10-8 × 0.5 / 10-6 ≈ 0.5%, so the expected multiplier is about 0.005 × 100M + 0.995 × 0.51M ≈ 1.0 million — close to the true value.
The paper also offers a second recipe for keeping zero out of the codebook. ZE-EDEN (Zero-Excluding EDEN calibration) keeps TorchAO's positive floor but rescales each 128-value block of v so small values are inflated less, which restores part of the large-multiplier tail that the floor cut off. Both recipes also change how the first moment m is stored. They use NF4, whose levels sit where most m values actually are, and they switch only the LM head's first moment to stochastic rounding for the final 10% of training. The authors added that switch after matched checkpoint experiments localized a late instability at 2.7B parameters to round-to-nearest on that one tensor's m; they offer a persistent inward bias, building up step after step, as a possible explanation. On a 1.4B-parameter GPT-style model, the validation-loss gap to 32-bit AdamW fell from +0.0521 with TorchAO 4-bit to +0.0156 with ZIP-SR — the paper's 70% headline — at the same storage cost. In full-parameter fine-tuning of Qwen3-8B-Base and Llama-3.2-3B, both recipes reached lower validation loss than TorchAO 4-bit, with downstream scores close to 32-bit AdamW.
| Recipe | Second moment v | First moment m | Gap to 32-bit, GPT-style 1.4B |
|---|---|---|---|
| TorchAO 4-bit AdamW | zero removed from codebook, round-to-nearest | 4-bit dynamic codebook | +0.0521, mean of 3 seeds (Table 4) |
| ZE-EDEN | zero removed, round-to-nearest, per-block EDEN rescale | NF4; LM head stochastic in last 10% | +0.0184, mean of 3 seeds (Table 4) |
| ZIP-SR | zero kept, stochastic rounding in preconditioner space | NF4; LM head stochastic in last 10% | +0.0156, mean of 3 seeds (Table 4) |
The paper's broader point: optimizer-state quantizers should account for how stored values affect later updates, not only for the rounding error in the stored number. For model weights the stored value and the used value are the same thing, so standard weight-quantization methods can minimize plain rounding error. AdamW's second moment is different: it passes through 1/(√v + ε) before it touches a weight, and that function is steepest near zero, exactly where the quantizer chooses between the zero level and the first positive level. Small error in the stored v can therefore coexist with large error in the next step's multiplier, which is worth keeping in mind when deciding what to quantize.
Goes deeper in: LLM Internals → Quantization → The Quantization Process
Related explainers
- TACO — one-sparse optimizer state — the other route to small optimizer memory: change the update rule so the state is sparse, instead of compressing AdamW's state
- UFP4 — E2M1 shrinkage bias — another case where round-to-nearest biases a 4-bit format and stochastic rounding is part of the fix