LLM·

DLFP in vLLM — Decode-latency feedback for chunked prefill — What does it mean?

The news. On September 29, 2026, Gaurav Agarwal, Ashish Garg and Isha Singhal posted Decode-Latency Feedback Prefill: A Model-Free Controller and Its Generalization Limits. They patched the vLLM V1 scheduler (vLLM 0.17.1) so that prefill chunks that share an iteration with active decodes are capped by a proportional controller, and evaluated it on Qwen3 models on A100 80 GB GPUs with open-loop Poisson arrivals and exact token accounting. The paper reports a clear win on Qwen3-0.6B on one GPU, and it reports, with equal weight, that the same mechanism fails on Qwen3-8B, Qwen3-32B and a two-GPU tensor-parallel setup. Read the paper →

Picture the ferry. Every trip carries the same commuters, and each commuter gets one stop closer to home per trip. When a truck of freight arrives, the clerk can load all of it onto the next trip, and then the trip is slow and every commuter on board waits for the freight. That is the interference problem: in continuous batching, one iteration mixes decode tokens from many users with prefill tokens for a new prompt, and the paper writes the iteration time as roughly decode work + prefill work + overhead. Every decoding request observes that whole time before it gets its next token, so a large prefill creates a stall that all of them feel at once.

The standard answer is to split the freight into crates: fixed-cap chunked prefill admits at most a set number of prompt tokens per iteration. But in that baseline the crate size is a hand-tuned setting. Big crates finish the prompt in fewer trips but make each trip slow; small crates keep trips fast but add trips, so the new user waits longer for the first token. The chunk size does not remove latency; it chooses who pays it — the new user through TTFT, or everyone already streaming through ITL. The paper's own fixed-cap controls show how fragile the hand-tuned setting is: a 2,048-token cap reduced SLO compliance to 287 of 300 requests, and a 4,096-token cap did not reliably clear the paper's 10% ITL-reduction gate.

DLFP replaces the hand-tuned crate with a clerk holding a stopwatch. It acts only when at least one request is already decoding — an isolated prefill, a ferry with no commuters, is never capped. After a guarded cycle it compares the observed interval with a target of 80 ms and sets the next cap to the current cap times target ÷ observed, rounded down to a multiple of 128 tokens and clamped between 1,024 and 6,144. A slow trip shrinks the next crate; a fast trip grows it. The rule needs no model of the GPU, the kernels or the batch — which is exactly why it is only as good as its stopwatch.

WITHOUT CHUNKINGWITH CHUNKINGt0t1t2t3t4t5t6t7t8t9Req APPDDDDReq BPPPPPPDD▲ Req A stalls (dark) while Req B prefills — TTFT increasest0t1t2t3t4t5t6t7t8t9Req APPDDDDDDDDReq BPPPPPPDD✓ Req A decodes every step — no stallingPrefill (P)Decode (D)Stalled

Here is the update the paper recorded in a live run, with the target fixed at 80 ms and the cap starting at its initial value of 3,072 tokens. The controller observed a 57.51 ms interval, faster than the target, so it multiplied: 3,072 × 80 ÷ 57.51 ≈ 4,273 tokens. Rounding down to a multiple of 128 gives 33 × 128 = 4,224 tokens — the exact value the paper logged, so the cap on the next prompt chunk grew by 37.5%. Run the same rule the other way (illustrative, same 80 ms target and 3,072-token starting cap): an observation of 120 ms gives 3,072 × 80 ÷ 120 = 2,048 tokens, and the cap shrinks by a third. On Qwen3-0.6B on one A100, this loop cut P99 inter-token latency in all three paired trials — for seed 801 from 122.22 ms to 91.95 ms — for a mean reduction of 27.7% (95% CI 21.0–34.3%), with all 76,800 paired token IDs identical. The price was on the other side of the trade: mean P99 TTFT rose 34.8%, still inside the declared SLO.

ConfigurationP99 ITL, baseline → DLFPOther effectSource
Qwen3-0.6B, 1× A100, 1.0 req/s−27.7% mean over 3 trialsP99 TTFT +34.8%; SLO passes 300/300§5
Qwen3-0.6B, 2 GPUs (tensor parallel)−12.7% mean, CI crosses zeroSLO passes 110 → 101 of 300§6.1
Qwen3-8B, 0.10 req/s19 ms → 281 msSLO passes 10 → 7 of 20§6.2
Qwen3-8B, 0.25 req/s1.763 s → 707 msP99 TTFT 12.33 → 14.81 s; SLO passes 2 → 0 of 20§6.2
Qwen3-32B, 2 GPUs, 0.05 req/s314 ms → 670 msSLO passes 5/12 both; energy per token +4.7%§6.3

Now the stopwatch. The prototype does not time the ferry from dock to far shore; it times how often the ticket booth calls the next boarding — the interval between scheduler calls on the CPU. vLLM runs with asynchronous scheduling, so the CPU can call the next boarding while the GPU still has queued work from the last one. On the small model the two clocks happened to agree. With larger kernels, a deeper asynchronous queue and tensor-parallel synchronization, they drift apart, and the controller can see a short host interval, conclude the trip was fast, and grow the crate while the device is still busy. The paper calls the small-model success an accidental correlation. A feedback loop with a proxy sensor can steer confidently in the wrong direction. The authors also tried a completion-aware revision that updated the controller when the scheduler output returned and bounded each step; on Qwen3-32B under synchronous scheduling it still lost (P99 ITL 41.1 ms baseline versus 194.8 ms), which they read as evidence against simple retuning. Their proposed next step is a controller timed on real completion in runtimes that expose it — a design the paper motivates but has not yet tested. This is the same boundary the scheduler's two halves draw in vLLM: the budget is spent in schedule(), but the truth about the iteration only arrives when the output comes back.

There is a second, quieter lesson in the 8B row. At 0.10 req/s there was almost no interference, so the baseline P99 ITL was 19 ms; DLFP cut prompts into more pieces anyway and P99 ITL rose to 281 ms. The authors note that splitting one large stall into several moderate stalls can increase the fraction of token gaps that reach the global P99. One long pause hurts one gap per user; five medium pauses can each land in the tail. That is why the paper judges every configuration on goodput — requests that met TTFT, mean ITL and end-to-end limits together — and not on ITL alone: at 0.25 req/s the 8B run moved latency from ITL into TTFT and SLO passes went from 2 to 0.

The authors see the opportunity in concurrent edge inference on phones and CPU-only laptops, which they have not yet evaluated: a laptop that generates an answer while it indexes a document, or several agent branches on one GPU, where full prefill/decode disaggregation is not available. A synchronous CPU runtime would expose real completion time directly, which removes the faulty stopwatch. The paper is explicit that it has not tested phones or CPU laptops and that adaptive chunk sizing itself is prior art; its claim is this specific model-free controller and where it stops working.

Goes deeper in: LLM Serving → Prefill/Decode Disaggregation → Chunked Prefill

Related explainers

Frequently Asked Questions

Check what you knowMap your AI & GPU knowledge across every track — free, role-based