LLM·

Restore long-context speculative acceptance with SharpDraft — Cardinality-aware query scaling — What does it mean?

The news. On October 4, 2026, researchers at UNIST posted SharpDraft (arXiv 2610.05106). They observe that drafters such as DFlash, PARD and EAGLE 3.1 start a long generation far faster than plain decoding, but their advantage shrinks as the output grows, and the attention mass each drafter puts on its 32 highest-scoring keys falls at the same time. SharpDraft rescales the drafter's queries by a closed-form function of the key count. Across AIME-26, GPQA-Diamond and LongGenBench Diary it reports 2.59–3.19× geometric-mean end-to-end speedups over target-only decoding, with no change to the drafter's weights or the target model. Read the paper →

Picture yourself at a party, trying to follow three friends. Early in the evening the room holds a handful of guests, and most of your hearing lands on the voices you care about. By midnight the room is packed. Your friends speak exactly as loudly as before, yet the share of your hearing they get has collapsed, because your attention is a fixed budget and every new guest takes a slice of it. A drafter faces the same room. Each new token adds one more key to its context, and softmax divides one fixed unit of attention across all of them.

The paper makes this precise with an exact identity. Split the visible keys into the top 32 and the rest. The log-odds of the attention on the top group (the log of its share divided by everyone else's share) equals a contrast term (the log of the top group's average exponentiated score minus the log of the rest's average) minus the log of how many competitors there are per top key, ln((N − 32) / 32). Hold the contrast fixed, double the number of competitors, and the attention odds on the top group halve. Figure 1 of the paper shows the drafters' top-32 attention mass falling as generation goes on, alongside their throughput; the authors note that this trend alone does not prove the key count is the cause, since content and positions change too. Their hypothesis is that dilution hurts a drafter in a specific way: DFlash's Qwen3-8B drafter, for example, was trained on sequences of at most 3,072 tokens, so a long reasoning trace takes it well past the lengths it learned on. A flatter drafter may guess less like the target, the target rejects more of the block, and the verification step accepts fewer tokens per pass.

The fix is the focus knob. Multiplying the query by a factor β above 1 multiplies every attention score by β, which stretches the gaps between scores without reordering them. A sharper query raises the contrast term, and the rule is sized so that this rise roughly offsets what the crowd took from the competitor term. It is an approximation: it does not guarantee that attention mass or acceptance is fully restored, and the top-scoring keys are not always the ones that matter for the task. SharpDraft sets the knob from the head count alone: no score inspection, no top-k selection, no training. The cached keys and values stay as they are, so nothing in the KV cache is rebuilt.

Draft Model (1B) — generates K=5 tokens
Paris.Itis
↓ verify all at once↓
Target Model (70B) — one forward pass
✓ Par✓ is✓ .✗ It→The— is
accepted
rejected → corrected
discarded
3 accepted + 1 corrected = 4 tokens from 2 forward passes
↻ repeat until done

In the paper's primary DFlash runs the rule has three fixed settings: a reference count of 2,048 visible keys (below it the scale stays at 1), k = 32, and a slope s = 5, with the scale capped at e0.5 ≈ 1.65; other drafters use other reference counts (PARD uses 8,192). The rule is β = min(e0.5, 1 + ln(r) / 5), where r = max((keys − 32) / (2,048 − 32), 1). Now walk one long context. At 16,416 visible keys, r = 16,384 / 2,016 ≈ 8.13. If the contrast stayed fixed, the crowd would cut the top group's attention odds by that factor of 8.13: a group that held 50% of the attention at 2,048 keys (odds 1 to 1) would drop to odds of about 1 to 8.13, or about 11% (10.96%) (the 50% starting share is illustrative; 2,048, 32 and 5 are the paper's settings). SharpDraft's scale at that point is β = 1 + ln(8.13) / 5 = 1 + 2.10 / 5 ≈ 1.42. The slope of 5 is the paper's stand-in for how fast contrast grows as β rises: if contrast gains about 5 log-odds units per unit of β, then raising β by 0.42 buys back about 2.1 units, which is ln(8.13), the exact amount the crowd took away. Under that fixed-slope assumption, the knob's 0.42 turn is sized to cancel the factor-of-8 crowd. The cap of 1.65 is reached near 51,700 visible keys, after which the scale stops growing.

DFlash on Qwen3-8BTrainable paramsPeak memorySpeedup vs plain decoding (geo-mean)Source
Frozen drafter026.87 GiB1.489×Table 1
Frozen + full adaptation1.049B41.29 GiB1.472×Table 1
Frozen + LoRA (rank 8)3.23M31.63 GiB2.144×Table 1
Frozen + SharpDraft026.87 GiB2.675×Table 1

The table shows why the knob beats retraining the listener. On the same drafter and target, query scaling lifted the geometric-mean speedup from 1.489× to 2.675× while adding zero trainable parameters and matching the frozen drafter's reported 26.87 GiB peak memory, whereas full adaptation added 14.42 GiB and ended slightly slower than doing nothing on this model. The paper reports the same direction on the other two drafter families: with EAGLE 3.1 on Qwen3-8B, mean accepted length rose from 3.152 to 3.302 on AIME-26 and from 5.668 to 5.968 on Diary.

The authors are careful about the limits, and the metaphor shows why. Turning up your focus makes the loudest voices louder; it cannot tell you that you were listening to the wrong friend. Because scaling never reorders scores, it fixes a drafter whose attention is too flat, not one whose attention points at the wrong tokens. The paper also finds that restoring the attention mass more exactly adds little further acceptance, that the slope and reference count should be re-validated for each new drafter, that the strongest end-to-end evidence is for DFlash under greedy decoding (always taking the single most likely token), and that behaviour under sampling, batching and concurrent serving is untested. So treat this as a cheap fix for one cause of the long-context slowdown, and consider it for any drafter you choose for a long-output workload.

Goes deeper in: LLM Serving → Speculative Decoding → Draft Model Variants

Related explainers

Frequently Asked Questions

Check what you knowMap your AI & GPU knowledge across every track — free, role-based