LLM·

RATIO paper — Token penalties for quantized overthinking — What does it mean?

The news. On September 30, 2026, a team from Shanghai Jiao Tong University posted RATIO (Reasoning Analysis and Token-level Inference Optimization) to arXiv. They quantized five reasoning models — three DeepSeek-R1-Distill-Qwen sizes (1.5B, 7B, 14B), DeepSeek-R1-Distill-Llama-8B and Qwen3-4B — to 3-bit weights with AWQ and GPTQ, and found that the quantized models both lost accuracy and wrote much longer chains of thought. RATIO reports up to 9.8 points of accuracy recovered and up to 51.3% shorter reasoning compared with the uncorrected quantized models, with no additional training. Read the paper →

Picture the understudy on opening night. The script in their hand is a blurry photocopy, so every few lines they stop: "Wait — was that the line?" "Hmm, let me check again." The play still ends, often with the right ending, but it runs far longer than when the lead actor performs it. A quantized reasoning model behaves the same way: rounding its weights to 3 bits does not only cost accuracy, it shifts probability toward hesitation tokens, so the model loops through extra verification before it answers. In the paper's examples, a 3-bit model reaches the correct answer to a simple word problem and then keeps writing "Wait, let me double-check that" and "Alternatively, perhaps" until it degenerates into repeated characters.

This matters because of how decoding works: every reasoning token is one more full pass through the model. Quantization makes each pass cheaper, mainly because each decode step reads every weight from memory and smaller weights mean fewer bytes to read (see why decode is memory-bound). But the bill is cost per token times the number of tokens, and overthinking raises the second factor. The KV cache, which these weight-only setups do not shrink, also keeps growing with every extra token (KV cache memory cost).

How RATIO writes the director's notes

The director does not hand the understudy a generic list of banned words. They watch both performers on the same scene and note which lines the understudy over-says, and by how much. RATIO does this in two stages.

  1. Find the tokens (Quantization-aware Reasoning Behavior Analysis, QRBA). Both models read the same fixed reasoning text through teacher forcing, so the comparison is on identical prefixes. For each candidate token, RATIO averages how much more probability the quantized model gives it than the BF16 model. It keeps only tokens whose probability rises under both AWQ and GPTQ, then checks them on the quantized model's own free-running answers: does the token appear more in failed and repetitive answers? A final context check removes tokens that do useful work such as genuine self-correction. For DeepSeek-R1-Distill-Qwen-1.5B this produced a set of 21 tokens; across the five models the sets include tokens such as "Wait", "Hmm", "But", "perhaps" and "something".
  2. Size each penalty (Token-Specific Penalty Determination, TSPD). For every position where the quantized model prefers a selected token more than the BF16 model does, RATIO computes the gap in logit space — how far that token's logit must move so its probability matches the full-precision one. It takes the median gap per token, keeps the smaller of the AWQ and GPTQ values, and divides by the median across the token set. At inference the only change is subtracting that per-token number from the token's logit before sampling — no weights change and no training step runs.

The per-token sizing is the difference from the earlier approach. A shared penalty treats "Wait" and "maybe" the same, even though quantization shifts them by different amounts and the set of over-used tokens differs from model to model; in the paper's tables, the normalized penalties range from about 0.36 to about 1.78.

Setting (DeepSeek-R1-Distill-Qwen-1.5B, avg of 5 benchmarks)AccuracyAvg CoT lengthSource
BF16 (full precision)60.13%10.09k tokensTable 1
AWQ-W3, no correction30.88%38.04k tokensTable 1
AWQ-W3 with one shared penalty on hand-picked markers37.79%25.86k tokensTable 1
AWQ-W3 with RATIO per-token penalties40.66%18.51k tokensTable 1

Where the decode budget actually goes

Extra reasoning tokens eat part of the per-step saving that quantization buys. Hold the model (DeepSeek-R1-Distill-Qwen-1.5B), the five-benchmark average and the AWQ 3-bit setting fixed. The BF16 model writes about 10.09k reasoning tokens per answer; the 3-bit model writes about 38.04k — 3.77× more tokens. Now take an upper bound on the per-token saving (illustrative): if every decode step were limited only by reading weights, going from 16 bits to 3 bits would make each step at most about 5.3× cheaper (16 ÷ 3, ignoring the per-group scales). Then the total decode work falls by only 5.3 ÷ 3.77 ≈ 1.4×, not 5.3×. With RATIO the 3-bit model writes 18.51k tokens, 1.83× the BF16 length, so the same upper bound gives about 5.3 ÷ 1.83 ≈ 2.9× — roughly twice the saving the uncorrected 3-bit model delivers, while average accuracy rises from 30.88% to 40.66%. Real step costs fall by less than 5.3× (attention, the KV cache and dequantization still cost time), which makes the token count matter even more.

What the results do and do not show

The headline numbers come from the smallest model. Gains shrink as models get larger: on the 7B, 8B and 14B models RATIO cuts reasoning length by roughly 4–22% and moves average accuracy by about −0.1 to +2.6 points, and on DeepSeek-R1-Distill-Qwen-14B with AWQ its accuracy is 0.07 points below the uncorrected model. Even at its best, the corrected 1.5B model (40.66%) stays well below the BF16 model (60.13%): in these experiments the penalties shorten the reasoning but do not close the accuracy gap to full precision. The main results use 3-bit weight-only quantization. The only 4-bit test is one appendix run on the 1.5B model with weights, activations and KV cache all at 4 bits, where RATIO raised accuracy from 45.27% to 48.48% and cut length by 12.9%; weight-only 4-bit, a common production setting, is not tested. The authors also note that each penalty is static, so a "Wait" that starts a useful correction is penalized the same as one that starts a loop.

Goes deeper in: LLM Internals → Quantization → Modern Methods

Related explainers

Frequently Asked Questions

Check what you knowMap your AI & GPU knowledge across every track — free, role-based