LLM·

OSFP4 tunes NVFP4 quantization — Joint smoothing and block-scale optimization — What does it mean?

The news. On October 6, 2026, researchers at the Hebrew University of Jerusalem posted OSFP4 (Optimized Smoothing and Scaling for NVFP4) on arXiv. For each linear layer it learns one scaling factor per input channel and the per-block scales together, against an approximation of the error of the layer's output, and it accounts for whether the weights will later be rounded plainly or with GPTQ. On Llama-3.1-8B-Instruct it reports the highest average accuracy among the 4-bit methods it compared, while keeping roughly 94–97% of vendor NVFP4 prefill throughput. The code ships as an add-on for LLM Compressor plus a vLLM plugin. Read the paper →

Picture the choir. Sixteen singers stand in one section, and the recording desk gives that section a single gain knob. That is NVFP4: every value is a 4-bit E2M1 number, and each block of 16 consecutive values shares one scale. E2M1 has only eight magnitudes — 0, 0.5, 1, 1.5, 2, 3, 4 and 6 — so the knob does most of the work. The default rule, absmax-to-6, turns the knob until the loudest singer sits exactly at the ceiling of 6, so nobody clips.

The trouble is everyone else. If one singer is far louder than the rest, the gain has to come down for that singer, and the quiet ones drop toward the noise floor. With absmax-to-6, any value more than 24 times smaller than its block's largest value is rounded to exactly zero — that is the outlier problem at the scale of sixteen numbers instead of a whole tensor. The paper adds a second, quieter point: even when no value is lost, several different knob settings keep the block in range, and the one absmax picks is often not the one with the smallest error.

Outliers force INT8 range to span –60 to 60 · most of the 256 grid slots fall on empty space

Weight value distribution-60-40-200204060
99.9%
0.1%
outliers
INT8 grid:-60 to 60 — 256 slots wasted on empty space
■ main distribution■ outliers■ INT8 grid

The obvious fix is to ask the loud singers to step back and the quiet ones to step closer. In a matrix multiply that is diagonal smoothing: scale input channel i of the activations by a factor, and divide column i of the weights by the same factor, so the exact product is unchanged — the playback undoes each step. This is the idea behind SmoothQuant and its relatives, but those were built for integer formats, where the rounding error grows with the largest value in the group, so balancing maxima is what matters.

For floating-point formats with a wide exponent range, the paper argues, this kind of smoothing is close to useless: a float's rounding error is roughly proportional to each value's own size, so scaling a channel up on one side and down on the other leaves the product's error about where it was. E2M1 breaks that argument. With only two exponent bits, a 4-bit float has too little range to ignore, so smoothing becomes useful again — but its job changes. It is no longer about balancing two maxima; it is about evening out the values within each block of 16, in the weight rows and the activation columns at once, so that each block's scale can do its best. You can see why the two precision formats behave so differently: FP8's E4M3 has range to spare, E2M1 does not.

So OSFP4 places the singers and sets the knobs together, on a rehearsal take. For each linear layer, using the fixed weights and a small set of calibration activations, it searches for one smoothing factor per input channel and the weight block scales at the same time (in the W4A4 setting, trial activation scales take part too). The target approximates the expected squared error of the layer's output — the matrix product — not the error of the weights or activations taken separately, and its formula changes depending on whether the weights will be rounded to nearest (RTN) or with GPTQ-style error feedback (SIC).

The hard part is that real rounding makes the error jump in steps as you turn a knob, which gives an optimizer nothing smooth to follow. The paper's trick is to analyze a randomized, multiplicatively dithered FP4 quantizer instead: its expected error is a smooth function of the smoothing factors and scales, and the paper relates it closely to the error of the real deterministic quantizer. That smooth stand-in, not the real rounding loss, is what the gradient steps minimize. After that first stage, the smoothing factors are kept and the weight scales are searched again on the actual E4M3 grid, with a criterion matched to the rounding procedure. The trial activation scales are thrown away: activations only exist at run time, so deployment still sets their scales with absmax-to-6.

Method (Llama-3.1-8B-Instruct, W4A4, absmax activation scales)What it changesMean accuracy, 4 tasksSource
BF16 (no quantization)reference79.22%Table 1
RTNabsmax scales, plain rounding75.65%Table 1
GPTQerror feedback across weight columns76.34%Table 1
NVIDIA released checkpointvendor-prepared; preparation not controlled by the authors76.11%Table 1
OSFP4 W-RTNjoint smoothing + scales, plain rounding76.36%Table 1
OSFP4 W-SICjoint smoothing + scales, GPTQ-style rounding77.03%Table 1

Two worked numbers make the mechanism concrete. First, one block (illustrative): say its loudest value is 12, so absmax-to-6 sets the scale to 2. A value of 0.4 becomes 0.2 on the E2M1 grid, below the halfway point to 0.5, and is stored as 0. A value of 0.6 becomes 0.3, rounds up to 0.5, and comes back as 1.0 — a 67% error. Now let smoothing halve the loud value's channel (and double that channel's weight column): the block's maximum becomes 6, the scale becomes 1, and 0.4 and 0.6 both come back as 0.5 — errors of 25% and 17%. The catch is the doubled weight column, which now has to survive its own 4-bit rounding; that trade is exactly why the factors and scales are tuned together.

Second, the paper's own numbers, holding the model (Llama-3.1-8B-Instruct), the format (W4A4) and the activation rule (absmax) fixed. Full precision averages 79.22% across MMLU-CoT, GSM8K, HellaSwag and WinoGrande; plain RTN drops to 75.65%, a 3.57-point loss. OSFP4 with GPTQ-style rounding reaches 77.03%, a 2.19-point loss — it wins back 1.38 of the 3.57 points, about 39% of what 4-bit rounding cost. WikiText-2 perplexity tells the same story: 7.88 for RTN, 7.77 for GPTQ-style rounding alone, 7.67 with the smoothing and scale optimization added, against 7.22 for BF16. The price, measured on an NVIDIA B200, is about 3–6% of prefill throughput, which the paper attributes primarily to the smoothing step before the attention-output and MLP down projections, where it cannot be folded into a preceding normalization layer.

Keep the size of the claim in view. The main comparison is one 8B model, with more models in the paper's appendix, and the gains are around a point of average accuracy — real, but not a closing of the gap to full precision. What makes the idea durable is the lesson underneath: with 4-bit floats, the shared block scale is the bottleneck, and the cheapest lever on it is how evenly you spread the values it has to cover.

Goes deeper in: LLM Internals → Quantization → The Outlier Problem

Related explainers

Frequently Asked Questions

Check what you knowMap your AI & GPU knowledge across every track — free, role-based