LLM·

AgSpec for coding agents — Retrieval-based speculative drafting — What does it mean?

The news. On October 1, 2026, Sumin Lee, Sukmin Cho, Suengjae Lim and Youngjin Kwon of KAIST posted AgSpec on arXiv. It is a layer on top of existing retrieval engines (the suffix automaton of SAM-Decoding and the suffix tree of SuffixDecoding), implemented in vLLM. On SWE-bench Verified and TeamBench with Devstral-24B, Gemma3-27B and Qwen3.6-27B, it compares against five retrieval drafters and EAGLE-3 and reports the highest or second-highest throughput in every evaluated setting. Read the paper →

Picture the secretary again. A good secretary does not wait for the lawyer to dictate every word of a letter that is mostly the same as last week's letter — they pre-type the familiar part from the binder and let the lawyer read it in one go. That is speculative decoding: normal decoding produces one token per sequential forward pass of the model, so anything that proposes several correct tokens at once and has the model confirm them in a single pass saves time. Most systems pre-type with a second, smaller model. A retrieval-based drafter pre-types by copying instead — it looks up the last few tokens in stored text and proposes what came after them there, with no model at all. That puts it among the draft model variants, at the cheapest end.

Coding agents are unusually good customers for copying, but only if the binder holds the right letters. A coder agent re-emits lines from a file it opened, an executor reproduces test logs, and a retry repeats most of the failed attempt. The paper argues that older drafters miss this for two reasons. First, availability: they index only the current request or a fixed datastore, so a file that the prompt truncated, or the previous attempt from an earlier call, is not in the binder at all. Second, matchability: the file is stored as it was read, but the agent writes it back as a unified-diff patch where every kept line starts with a space and every removed line with -. The text is the same; the tokens are not, so the lookup fails.

AgSpec fixes the binder with three corpora keyed by where the text came from. The session corpus holds the text of the current session — prompts, tool outputs and generated tokens — shared across its calls and agents. The workspace corpus holds each repository file the session opens — once as is, and again in the agent's emission form: a copy with every line prefixed by a space and another with every line prefixed by -, so any line the coder copies into its patch has a matching entry. The global corpus holds logs from past sessions on other tasks. Matches from the two live corpora win on length; the global corpus must beat them by a margin, because a long match in unrelated text is more likely to be a coincidence.

Draft Model (1B) — generates K=5 tokens
Paris.Itis
↓ verify all at once↓
Target Model (70B) — one forward pass
✓ Par✓ is✓ .✗ It→The— is
accepted
rejected → corrected
discarded
3 accepted + 1 corrected = 4 tokens from 2 forward passes
↻ repeat until done

The second fix is how far ahead to pre-type. A drafted token saves a sequential step if the model accepts it, but costs verification compute if it is rejected — and at larger batch sizes verification becomes compute-bound, so rejected tokens stop being nearly free. AgSpec drafts the match length times a per-agent scale that verification feedback moves up or down, and never more than a per-agent cap profiled offline. The cap comes from replaying past trajectories and finding the last draft position whose acceptance probability is still above the break-even threshold λ. The scale starts at 1.0 (draft as many tokens as were matched), is updated after verification steps (except when a fully accepted draft hits the cap, which says nothing about a longer one), and persists across sessions on the same server, so a coder agent that is copying long stretches drafts long, and a localizer that names short paths drafts short — and both adjust as a session drifts from turn to turn.

Where the cap comes from, with the paper's own illustrative example. Hold the break-even threshold at λ = 0.1: verifying one more draft position costs as much time as producing 0.1 tokens. Suppose an agent's 11th draft token is accepted with probability 0.12 and its 12th with 0.09. The 11th position returns 0.12 tokens in expectation for a cost of 0.1 — a small win, so it stays. The 12th returns only 0.09 for the same 0.1 — a small loss, so it goes. The cap for that agent is therefore 11. Change the agent and the probabilities change, so the cap changes with it. This is why a single global draft length is the wrong knob for a pipeline of five different agents: the same 0.1 cost line cuts each agent's acceptance curve at a different position.

DrafterWhat its binder holdsHow far it pre-types
PLD (prompt lookup)The current prompt and generated prefixFixed cap
RESTA prebuilt text or code datastoreFixed tree length
SAM-DecodingPer-request suffix automaton plus a static datastoreFixed cap
SuffixDecodingPer-request suffix tree plus a global tree of past responsesScales with match length by a preset factor
FastCoderCommon code plus repository files and a cache of verified sequencesFixed cap
AgSpecSession, workspace (also in diff form) and global corporaMatch length times a per-agent online scale from verification feedback, limited by a per-agent offline cap

AgSpec gets its gains by changing the binder and the length policy around existing retrieval engines, not the engines themselves. Across the main settings AgSpec reaches 2.27–4.37× the throughput of autoregressive decoding at batch size 1 and 1.08–4.76× at batch size 16, and on average 18.0% more throughput than the fastest prior method. Put on top of SAM-Decoding, it raises Gemma3-27B's speedup on SWE-bench Verified at batch size 16 from 2.72× to 3.61×. On LiveCodeBench, which has no repository, the faster AgSpec variant still reaches 4.87–7.94× autoregressive throughput, from the session and global corpora alone. For an agent's cost profile, that is decode time removed without changing a single output token. Note the contrast with prefix caching: a prefix cache reuses computation only for an identical leading prefix, while a retrieval drafter reuses text from anywhere — the middle of a log, a file opened three calls ago — and still lets the target model decide.

Goes deeper in: LLM Serving → Speculative Decoding → Draft Model Variants

Frequently Asked Questions

Check what you knowMap your AI & GPU knowledge across every track — free, role-based