LLM·

Reparameterize output heads before W4 quantization — Softmax-preserving head shift — What does it mean?

The news. On September 25, 2026, researchers at Adobe posted Softmax Reparameterization for Output-Head Quantization (arXiv 2609.31291). It quantizes only the output head of seven small and multilingual models (Gemma 3 and 4, Qwen3.5, Phi-4-mini, BLOOM, BLOOMZ, XGLM) to 4 bits, and shows that picking an equivalent version of the head before rounding recovers much of the lost accuracy. On Phi-4-mini, the packed 4-bit head cuts batch-one generation latency by 10.8% versus the 16-bit head. Read the paper →

Picture a map with one peak per word in the vocabulary. The height of each peak is that word's logit, and the reader at the end, softmax, only ever asks how much taller one peak is than another. So you can move sea level up or down as far as you like: every peak changes height by the same amount, the gaps stay the same, and the probabilities do not move at all. In numbers, logits [3, 2, 0] and [-2, -3, -5] give exactly the same softmax, about [0.70, 0.26, 0.04] (illustrative).

In the real model, moving sea level means editing the weights, not the logits. Each logit is one row of the output head multiplied by the hidden vector h. If you subtract the same vector v from every row, every logit drops by the same amount v·h, so the model's predictions are exactly unchanged for every possible input. The head therefore has a whole family of equivalent versions, and the trained one is just one member of that family. The paper walks along one line through that family: subtract c times the average row, for a scalar c. c = 0 is the original head and c = 1 is ordinary mean-centering.

Quantization sensitivity by layer positionEmbeddingFP16errors affect every tokenLayer 1 AttnFP16Layer 1 FFNINT8Layer 2 AttnINT8Layer 2 FFNINT4···INT4middle layers — most robustLayer N FFNINT4Layer N AttnINT8Output projFP16errors affect predictions
■ high sensitivity (FP16)■ medium (INT8)■ low (INT4)

The diagram above is the usual rule from What to Quantize: the final projection to vocabulary logits is kept in 16-bit, because no later layer can absorb its error. That rule is expensive. The head is a plain matrix-vector product over the whole vocabulary for every generated token, and single-token decode is memory-bound, so the cost is the bytes you read. The paper counts roughly 0.4–0.7 billion weights in the output heads of recent 3–4B models, about 10–17% of the model.

Now bring in the 16-mark ruler. W4 quantization puts each group of 128 weights under one shared scale, and the chosen range and scale decide where each weight lands on a grid of at most 16 levels. Moving sea level changes every weight in every row, so it changes those ranges and which values land near a mark. Equivalent heads in full precision become different heads after rounding, and some of them round much better than others.

The obvious choice, full mean-centering, makes the weights as small as possible, but the paper finds that is not the same as keeping predictions right. On Phi-4-mini, mean-centering produced the smallest raw logit error among the three coefficients the authors measured, yet a different c gave lower KL divergence. The reason is that softmax does not weigh all errors equally: an error shared by every token costs nothing, and an error on a token with near-zero probability costs little. So the method tries a 14-point grid of c values (including 0 and 1) with the real quantizer and keeps the one with the lowest validation KL against the original model. The chosen shift is folded into the weights before packing, so inference runs the same 4-bit kernel with no extra operation.

Phi-4-mini output headKL to BF16WikiText perplexityBatch-1 latency, 64 tokensSource
BF16 (not quantized)0.0011.651136.6 msTable 3
W4 AW-MSE0.9731.651014.8 msTable 3
W4 AW-MSE + shift0.2815.401013.4 msTable 3
W4 GPTQ0.1513.67not measuredTable 3
W4 GPTQ + shift0.0612.23not measuredTable 3

Worked example. Hold three things fixed: Phi-4-mini, a vocabulary of about 200K rows, and batch-one decoding. In BF16 the head alone is about 1.23 GB of weight reads per generated token (the paper's analytical figure). At 4 bits that is about a quarter, ~0.31 GB, before group scales. On an A10G GPU through vLLM, 64 greedy tokens take 1136.6 ms with the BF16 head and 1013.4 ms with the shifted W4 head: 10.8% faster end to end, from quantizing one matrix. Plain W4 already gets that speed, but it pays for it: perplexity goes from 11.65 to 31.65. The shift keeps the speed and brings perplexity back to 15.40. With GPTQ as the base quantizer the shift still helps, taking KL from 0.15 to 0.06, so it stacks on top of stronger quantizers instead of replacing them.

The method has limits the paper states plainly. At W4, several heads already round well and gain little; the big wins are on heads with large baseline error, like Phi-4-mini, BLOOM and XGLM. Models that squash logits with a tanh soft-cap need the rank-one correction, and models that tie the output head to the input embedding must store a separate quantized copy of the head.

Goes deeper in: LLM Internals → Quantization → What to Quantize

Related explainers

Frequently Asked Questions

Check what you knowMap your AI & GPU knowledge across every track — free, role-based