Minimize speculative-decoding rounds directly with EDR — Expected Decoding Rounds vs block-local losses — What does it mean?
The news. On October 7, 2026, Yunxiao Zhao and Changxiao Cai of the University of Michigan posted Training Parallel Speculative Draft Models by Directly Minimizing Expected Decoding Rounds on arXiv. They model speculative decoding as a Markov reward process, derive the EDR objective and an exact gradient for it, and add an offline evaluator that counts expected rounds from target-model outputs without running speculative decoding. One epoch of EDR fine-tuning raised mean accepted length for the DSpark and DFly drafters on all nine benchmarks they report. Read the paper →
Picture a multi-day hike where you plan each day's whole route from the tent, without seeing the trail ahead. A ranger walks your plan from the start, step by step. At the first wrong turn the day ends: the ranger puts you one step onto the right path, you camp there, and tomorrow you plan again from that new campsite. The trip's length in days depends on two things at once: how good each plan is, and where yesterday's first mistake left your tent.
That is speculative decoding with a parallel drafter. Each round, the drafter proposes a whole block of tokens in one forward pass, conditioned on the prefix committed when the round began. The target model then checks the block left to right: each draft token is accepted with a probability set by the ratio of the target's to the draft's probability for that token, the first rejection ends the round with a corrected token from the target, and the rest of the block is discarded.
For an autoregressive (AR) drafter the campsite does not matter. An AR drafter sees the full prefix before every guess, so its guess for position 50 is the same whether the round started at position 45 or 48, and the paper notes that matching the target's next-token distribution at every position is already aligned with fewer rounds. A parallel drafter's guess for position 50 depends on where the round started: from position 45 it predicts five positions ahead, from position 48 only two. The same output position gets different draft distributions, and which one is used depends on earlier acceptances that the drafter itself caused. Draft model variants shows where parallel drafters sit among the drafter designs.
EDR makes the campsite explicit. Fix one output that the target produces. A state records where the current round started and which position is being checked. From each state, an acceptance moves one step along the same round, and a rejection costs one round and starts a new round after that position. The expected number of rounds is one for the first round plus, summed over every state, how likely decoding is to visit it (its occupancy) times the chance that a new round starts there. A new round starts after a rejection, or after a fully accepted block and its bonus token; an end-of-sequence token ends decoding instead. The paper proves this count is exact for a fixed output, then averages the rejection term over the target's full distribution to get a local cost: the total-variation distance (a measure of how much two probability distributions disagree) between draft and target at that state, with the end-of-sequence token excluded and the result normalized by the target's remaining probability mass. Averaged over target outputs, this EDR objective equals the expected number of rounds. The objective has no extra hyperparameters, and an exact TD-form gradient trains it from target-model rollouts.
| Training objective | What it scores | Sees the effect on later rounds? |
|---|---|---|
| Per-token distribution matching (CE / KL) | Draft vs target distribution at each position | No |
| Block-local, e.g. E2E | A surrogate for expected tokens accepted within one round | No |
| EDR | Occupancy-weighted cost of starting a new round; averaged over target outputs, equals expected rounds | Yes |
Why not simply maximize the tokens accepted in each round? Back on the trail: a plan that gets you a little farther today can leave your tent at the foot of a confusing junction, where tomorrow's plan is likely to go wrong early. The paper proves that a drafter that maximizes expected accepted length in every round can still be a constant factor worse on overall MAL than the best drafter of the same class, because the per-round optimum ignores where it leaves the next round's start. The proof uses a constructed target distribution, so it shows that the gap can exist, not how large it is for real models.
The same formulation gives a second tool: an exact offline evaluator. Because verification produces the same output distribution for every drafter, you can sample target-model outputs once and compute each candidate drafter's expected rounds on those same outputs, with no speculative-decoding run. Comparing two drafters on identical trajectories is a paired comparison, which removes the noise of separately sampled evaluation outputs; the paper computes all of its reported MAL numbers this way.
| Target + drafter | Benchmark | Original MAL | E2E MAL | EDR MAL | Source |
|---|---|---|---|---|---|
| Qwen3-4B + DSpark | AIME25 | 5.46 | 5.57 | 5.59 | Table 2 |
| Qwen3-4B + DSpark | Arena-Hard | 3.59 | 3.67 | 3.70 | Table 2 |
| Qwen3-8B + DFly | AIME25 | 5.15 | 5.18 | 5.22 | Table 2 |
| Qwen3-8B + DFly | Arena-Hard | 3.18 | 3.21 | 3.25 | Table 2 |
Here is what those numbers mean in rounds. Hold three things fixed: the target (Qwen3-4B), the drafter architecture (DSpark), and the answer length (an illustrative 1,000 tokens), and use rounds ≈ answer length ÷ MAL. On AIME25 the original DSpark has MAL 5.46, so the answer takes about 1,000 ÷ 5.46 ≈ 183 rounds. After one epoch of EDR the MAL is 5.59, about 1,000 ÷ 5.59 ≈ 179 rounds: roughly 4 fewer target passes, about 2.3%. On Arena-Hard, 3.59 → 3.70 moves the answer from about 279 to 270 rounds, about 8 fewer passes (≈3%). Against the block-local E2E fine-tune, the fairer baseline, the margin is smaller: 5.57 → 5.59 on AIME25 is less than one round per 1,000 tokens. Because EDR changes only training, the drafting and verification work per round stays the same, so fewer rounds should mean lower latency, but the paper reports MAL, not wall-clock speedup.
EDR costs more training compute than block-local losses, because it computes occupancy weights and value functions for many possible round starts (anchor positions sampled from each trajectory). The experiments cap every method at 512 anchors per trajectory and pick EDR's anchors by importance sampling. The framework also assumes a fixed draft block size, and extending it to adaptive block sizes is listed as future work.
Goes deeper in: LLM Serving → Speculative Decoding → Draft Model Variants
Related explainers
- PPOW: window-level RL for speculative drafters — another way to train a drafter on speed instead of per-token agreement, for autoregressive drafters
- BlockPilot: instance-adaptive draft block sizing — the adaptive block size that EDR's framework does not yet cover
- Nucleus acceptance vs exact verification — what changes when verification stops being exact, which EDR's offline evaluator relies on