LLM·

TRACE FP4 RL — Rollout-guided FP4 rounding — What does it mean?

The news. On October 6, 2026, researchers from Alibaba Group and Ohio State University posted TRACE (Train-Rollout Quantization Alignment via Compact Guidance), an FP4 quantization framework for reinforcement-learning post-training of Mixture-of-Experts models. It runs rollouts with FP4 weights, activations and KV cache, and keeps the training step aligned with them. On Qwen3.5-35B-A3B, TRACE averaged 75.3 across four reasoning benchmarks, against 74.9 for BF16 rollout and 59.6 for plain FP4 QAT. The paper evaluates four MoE models, from 35B to 2.4T total parameters. Read the paper →

Picture the two cashiers again. Each cashier is accurate on its own — both round to the nearest coin — yet they disagree, because their two prices sit on opposite sides of the halfway line. That is the situation in RL post-training with low-precision rollout. The rollout engine generates trajectories with the policy in FP4; the training engine then recomputes the probabilities of the same tokens to build the gradient. Their high-precision activations already differ slightly (different kernels, different batch shapes, sometimes a policy a few updates newer), and rounding to a coarse grid can turn that tiny difference into a full step on the grid. The paper measures that vanilla NVFP4 adds substantial train–rollout discrepancy, on top of the BF16 difference, across the model's layers.

Earlier fixes tried to make each cashier more accurate. Methods such as QUADS and Rollout-ResQ reduce each path's quantization error relative to its own BF16 value. But making each path closer to the truth does not make the two paths closer to each other. The paper gives a counter-example: QUADS keeps the training activation in BF16 and adds a residual correction on the rollout side, which lowers the rollout error but raises the gap between the paths in a case where plain FP4 had rounded both values to the same codeword. In RL that gap matters more than the absolute error, because the update weighs each token by the ratio of its training-side to rollout-side probability. In a Mixture-of-Experts model a small drift can also move a router score enough to send a token to a different expert, which amplifies the mismatch.

32-bit float — virtually continuous

01234π = 3.14159Quantized: ≈ 3.1416Δ = 0.00000Representable values (0 → 4)

TRACE changes what the training cashier rounds toward. During rollout it records which FP4 codeword each captured value became; during training it chooses, between the two neighboring codewords, the one closest to the rollout's choice instead of plain round-to-nearest. It does this for the activations of the routed experts and for the FP4 KV cache; the other modules stay in BF16. With the exact rollout codeword, the same rollout-side scale and the same two candidates, this by construction never makes the local gap at a captured site larger than round-to-nearest would; the default TRACE works from a compressed record (next paragraph), so it approximates that reference. The paper is explicit about the limit: even the full-information guarantee does not promise that the whole network's output gap shrinks, and it does not remove the real difference when the training policy is newer than the rollout policy. It reduces the extra amplification that inconsistent rounding adds.

The catch is the size of the sticky notes. For Qwen3.5-35B-A3B with responses up to 256K tokens, the paper estimates full guidance at about 50 KB per generated token; with 4,096 trajectories in one RL step that is up to ~51 TB per step, roughly three hours of writing at a 5 GB/s storage bandwidth. Two observations shrink it. In over 99% of the mismatched values, training and rollout land on adjacent FP4 codewords, so a small amount of mantissa information plus the scale is usually enough to identify the rollout's neighbor. The corrections also cluster in the deeper layers, so the default TRACE keeps only mantissa and scale information from the latter half of the layers and reconstructs an approximate rollout reference from it. In the paper's end-to-end measurement the RL step time rises from 664 to 713, a 7.4% overhead over plain FP4 rollout.

Worked example (illustrative numbers). Hold three things fixed: both engines use the same block scale, the two E2M1 codewords around the value are 2 and 3, and the halfway line is therefore 2.5. After scaling, the rollout engine sees 2.51 and the training engine sees 2.49 for the same token at the same site, a BF16 gap of 0.02. Round-to-nearest sends rollout to 3 and training to 2, so the quantized gap is 1.0: a 50× amplification from one rounding step. With the rollout's note, TRACE sees that the training value's neighbors are 2 and 3, that rollout chose 3, and picks 3: the quantized gap is 0. This shows the amplification TRACE targets at each captured routed-expert activation and KV entry; it illustrates the mechanism, not a guaranteed zero gap at every site.

MethodWhat it tries to minimizeAvg. score, Qwen3.5-35B-A3B (4 reasoning benchmarks)
BF16 rolloutNo FP4 in rollout (reference)74.9 (paper Table 1)
FP4 QATTraining-side error, by simulating FP4 in the forward pass59.6 (paper Table 1)
QUADSEach path's error against its own BF16 value68.8 (paper Table 1)
TRACEThe gap between the two FP4 paths75.3 (paper Table 1)

One more result changes how to read the FP4 cost. On HMMT25, TRACE starts below BF16 rollout because of the quantization error, then catches up after about 120 RL steps and stays comparable. Once the gap between the two paths is controlled, the policy can adapt to its own FP4 execution during RL, instead of only losing accuracy to it. The paper reports the same pattern on larger models: 70.6 vs 68.8 for BF16 on Terminal-Bench (Qwen3.8-Flash-Next) and 90.2 vs 90.3 on GDPval (Qwen3.8-2.4T-A95B).

Goes deeper in: LLM Internals → Quantization → The Quantization Process

Related explainers

Frequently Asked Questions

Check what you knowMap your AI & GPU knowledge across every track — free, role-based