LLM·

Relax speculative verification with target-model nucleus acceptance — Nucleus acceptance vs exact verification — What does it mean?

The news. On October 6, 2026, Shuhao Li, Fanghua Ye, Wanyu Lin, Tianyu Yuan and Xiaoyu Shen posted Nucleus Speculative Decoding: Plausibility-Aware Verification Beyond Exact Distribution (arXiv 2610.07822). They add one acceptance condition to speculative decoding — keep a draft token if it is in the target's top-p nucleus — prove that the extra acceptance equals the output's deviation from the target, and report throughput up to 5.16× over plain autoregressive decoding and up to 3.15× over standard speculative decoding, on math and code benchmarks with Nemotron and Qwen3 models of 3B to 14B parameters. Read the paper →

Picture the door of a club. A promoter walks up with a line of five guests, and the host has an ideal crowd mix in mind. The bouncer works to a quota: if the promoter brings more of some type of guest than the host wants, the surplus is turned away at random, even when the host likes that guest. And the first guest turned away ends the night for everyone behind them in the line.

That is how standard verification works. The draft model proposes a block of tokens, the target model scores all of them in one pass, and each draft token x is kept with probability min(1, p(x)/q(x)). Verification stops at the first rejection, the target resamples one replacement token, and every later draft token is discarded — however good it was.

The quota exists for a good reason: the acceptance rule plus the replacement draw make the output distribution exactly the target's, so standard speculative decoding changes speed, not answers. The cost is that the rule asks a narrow question. It compares the draft's probability with the target's; it never asks whether the token is a plausible choice for the target. When the draft is over-confident (q > p) about a token the target also ranks highly, that token can still be rejected.

The paper measures how often this happens. 81.5% of rejected tokens were in the target's top 10, and 52.6% were inside its top-p 0.95 nucleus. Draft models also tend to be more sure of themselves than their targets: where NSD's extra acceptances occur, mean draft entropy (a measure of how spread out the probabilities are; lower means more confident) is 0.865 against 1.101 for the target. That over-confidence is exactly what triggers these rejections.

NSD gives the bouncer a guest list. At each position the target's probabilities are sorted, and top tokens are added until they cover θ = 0.95 of the probability mass — the same top-p nucleus that nucleus sampling draws from. A draft token is kept if it passes the quota or if it is on the guest list. Nothing else changes. The target's probabilities are already computed by the verification pass, the replacement draw and the bonus token stay the same, and the method needs no training. Because it touches only the acceptance check, it works with any drafter: the paper tests linear self-speculation on Nemotron diffusion models and the DSpark drafter on Qwen3 in its main results, and DFlash, EAGLE-3 and multi-token prediction in its appendix.

Draft Model (1B) — generates K=5 tokens
Paris.Itis
↓ verify all at once↓
Target Model (70B) — one forward pass
✓ Par✓ is✓ .✗ It→The— is
accepted
rejected → corrected
discarded
3 accepted + 1 corrected = 4 tokens from 2 forward passes
↻ repeat until done

The price is exact and simple to state: the extra acceptance is the deviation. Call R the probability that standard verification rejects a proposal, and ES the part of the draft's surplus probability that lies inside the nucleus. Standard verification keeps a proposal with probability 1 − R; NSD keeps it with probability 1 − R + ES. The paper proves that, at a single position — with proposals drawn from the draft and the standard replacement draw on rejection — the total variation distance between NSD's next-token distribution and the target's is exactly ES. In club terms, every extra admit shifts the room's mix by the same amount, and the shift can never exceed R. Over a sequence the per-token errors compound: if every step's actual output deviates by at most ε, a coupling argument (running both decoders on shared randomness) shows they produce the identical T-token sequence with probability at least (1 − ε)T. The authors note that applying this to a specific speculative backend means first checking that its proposal, recovery and bonus steps meet those assumptions.

Here is how speed and fidelity pull on the same number (illustrative). Hold the draft block at K = 7 tokens, the block size the paper pairs with Qwen3, and treat each position as accepted independently. If standard verification accepts 60% of proposals, a pass emits on average (1 − 0.68) / (1 − 0.6) ≈ 2.46 tokens. If 20 points of the rejected mass lie inside the nucleus, NSD accepts 80%, and a pass emits (1 − 0.88) / (1 − 0.8) ≈ 4.16 tokens — about 1.7× fewer target passes for the same answer length — a pass count under these assumptions, not a measured speedup. But by the result above, that same 0.20 is also the per-token deviation, far too large for the bound to promise anything. The paper's own illustration uses a much smaller error: with ε = 0.001 per token and T = 32 tokens, the output matches the target with probability at least 0.99932 ≈ 96.85%. So the bound certifies closeness only when the extra admits per token are small; at acceptance gains as large as this example, it says little, and the evidence for quality has to come from benchmarks.

Model and drafterτ, standard SDτ, NSDNSD speedup vs autoregressiveSource
Nemotron-Labs-Diffusion-8B, linear self-speculation, 32-token blocks4.9811.824.35×paper, Table 1
Nemotron-Labs-Diffusion-14B, same setup4.6711.544.36×paper, Table 1
Qwen3-8B, DSpark drafter, 7-token blocks5.416.694.34×paper, Table 1
Qwen3-4B, DSpark drafter, 7-token blocks5.366.604.00× (3.13× under standard SD)paper, Table 1

The gain grows with block length: a 32-token block has much more to lose from one early rejection, and on the 8B and 14B Nemotron models the guest list more than doubles τ, against roughly +1.2 to +1.3 tokens on 7-token Qwen3 blocks. The two setups also use different drafters, so this is not a controlled comparison. All figures are overall averages across the paper's benchmarks at temperature 1.0 and θ = 0.95.

The quality result is mixed rather than free. Across 30 model-benchmark pairs, NSD beat standard speculative decoding on accuracy in 15, tied in 1 and fell behind in 14, with differences from −1.58 to +4.14 percentage points and +0.40 points on average. Qwen3-4B lost about one point overall (68.65% to 67.66%). The threshold θ is a real dial, and not a one-directional one: on Nemotron-14B, lowering θ from 0.95 to 0.85 raised accuracy on GSM8K, a set of math word problems, from 92.87% to 93.86% and cut speedup (5.16× to 4.86×), while on Nemotron-8B lowering it to 0.90 raised both on GSM8K but lowered the pass rate on MBPP, a Python coding benchmark. The tested models are 14B parameters or smaller, a limit the authors name themselves. The practical reading: lossless speculative decoding stays the safe default, and relaxed acceptance is a tradeoff you tune and evaluate per task, the way you would evaluate a quantized model.

Goes deeper in: LLM Serving → Speculative Decoding → The Verification Algorithm

How top-p picks a variable-size set of candidate tokens is covered in LLM Internals → Text Generation → Top-K and Top-P Sampling.

Related explainers

Frequently Asked Questions

Check what you knowMap your AI & GPU knowledge across every track — free, role-based