JIL attack — Adversarial length suffix vs predicted-length SJF scheduling — What does it mean?
The news. On October 2, 2026, Yuyang Dai, Rana Shahout and Mahmood Sharif posted Jumping the Line: Exploiting Length Predictions in LLM Scheduling to arXiv. Using the TRAIL scheduler on vLLM as a case study, they optimize a 20-token suffix per request against the scheduler's length probe. They report predicted lengths cut by up to 83.4% across four models (Llama-3-8B, Qwen2.5-7B, Mistral-7B, Llama-3.1-70B) and attacked requests finishing up to 1.53× faster on average in end-to-end serving runs. They also test two scheduler-side defenses: a prediction floor and coarse prediction bands. Read the paper →
Picture a supermarket with one register and a clerk who wants the line to move fast. The clerk lets whoever has the smallest cart go next, and judges cart size by a quick glance. That works well when people are honest: a shopper with three items no longer waits behind a full weekly shop. It is the same reason an LLM engine might leave plain arrival order behind. A first-come, first-served waiting queue is fair but slow on average, because one 2,000-token essay at the front holds up every one-line answer behind it.
The weak point is the glance. The clerk never counts the items; the clerk only judges how the cart looks. Throw a sheet over a full cart and it looks small, so it goes to the front. The items under the sheet are still rung up one by one, so the register is busy just as long, and every honest shopper who was moved back waits longer.
In the LLM version, the glance is a probe on the model's hidden state, and the sheet is about 20 optimized tokens. The server runs the prompt through the model to layer 11, averages the per-token vectors, and gives that summary to the probe. The probe outputs probabilities over ten length bins, and the predicted length is the probability-weighted average of the bin lengths. The scheduler sorts on that number: a smaller prediction means an earlier slot. vLLM's own priority policy already keeps its waiting queue as a heap ordered by a priority key, so a predicted-length scheduler is, in effect, computing that key from the prompt.
JIL attacks the key, not the model. It uses GCG to choose suffix tokens that minimize the probe's predicted length, with gradients flowing back through the probe and the first 11 layers. The attacker needs white-box access to the model and the probe while building the suffix, and builds one suffix per request. Resource-consumption attacks, such as sponge examples, try to make a request more expensive. JIL leaves the work mostly as it is and changes only the label on it. In the paper, the drop in predicted length is much larger than the drop in actual length: attacked open-ended answers stay several hundred tokens long on average.
Two costs come with the trick, and only one of them falls on the attacker. Honest requests pay first: in the paired serving runs (Llama-3-8B on one B200 GPU, 25 requests per workload), attacked requests moved forward in completion order and benign requests moved back, and with only 16 sequence slots the benign 95th-percentile slowdown reached 1.45×. The attacker pays in quality, which differs by model: on Alpaca, Llama-3-8B's win rate (how often a judge model preferred its answer over the dataset's reference answer) fell from 84.0% to 23.0% while its mean answer length stayed nearly the same. The headline speedup also needs a caveat. Attacked requests generated somewhat fewer tokens, so the paper notes that the measured speedups combine better queue position with shorter answers.
| Defense setting (16 slots) | Attacker speedup | Attacker position gain | Benign tail slowdown | Benign requests slowed |
|---|---|---|---|---|
| No defense, grouping batch (Table 3) | 1.42× | 15.25 pts | 1.43× | 55.3% |
| Round up to 32-token bands (Table 3) | 1.42× | 15.68 pts | 1.38× | 50.7% |
| Round up to 64-token bands (Table 3) | 1.35× | 12.96 pts | 1.21× | 40.3% |
| Round up to 128-token bands (Table 3) | 1.25× | 8.80 pts | 1.13× | 44.0% |
| No defense, floor batch (Table 3) | 1.43× | 14.24 pts | 1.35× | 37.0% |
| Floor at 80 tokens (Table 3) | 1.36× | 12.96 pts | 1.26× | 27.0% |
| Floor at 120 tokens (Table 3) | 1.20× | 10.88 pts | 1.22× | 59.0% |
The size-band defense gave the smallest honest-request tail: stop the scheduler from seeing fine differences. The paper rounds every prediction up to a multiple of a band width q (at least one band). Hold three things fixed: q = 128 tokens, 16 slots, and three illustrative predictions. An honest short request is predicted at 100 tokens, an honest long one at 300, and the attacker's 300-token request is pushed down to 66 by its suffix (a 78% cut, inside the paper's "up to 83.4%"). Without bands, 66 beats 100, so the fake request goes ahead of an honestly short one. With bands, 66 and 100 both round to 128 and the honest 300 rounds to 384. The suffix still moves the attacker down two bands, ahead of honest long work, but it now ties genuinely short work instead of passing it; which of the two goes first inside a band depends on the scheduler. The bands reduce the advantage; they do not remove it. In the measured runs this cut the attacker's speedup from 1.42× to 1.25× and the benign tail slowdown from 1.43× to 1.13×. For an honest request in that tail that took 10 s in the clean run (illustrative), that means about 14.3 s without bands and 11.3 s with them. At 64 slots, where few requests queue, the bands changed the attacker's speedup very little (1.27× to 1.25×).
Read the table as a trade-off, not a fix. A prediction floor at 120 tokens gives the smallest attacker speedup, but it slows the largest share of honest requests (59.0%); the paper reports this pattern without establishing its cause. The general lesson for SLO design is that any prediction that decides who gets shared compute is a security boundary. The paper suggests further defenses but does not test them: adversarial training of the probe, checking whether a prediction stays stable across small edits of the prompt, and adding waiting time to the priority so that a request moved back cannot wait forever.
Goes deeper in: LLM Serving → Inference Engine → The Scheduler
Related explainers
- CodeSpear strips an LLM's ability to refuse — Grammar-constrained decoding jailbreak — another attack that targets a serving component instead of the model's weights.
- Denoising Surface paper — Denoising Workload Surface — the same shortest-job-first idea, with cost prediction for diffusion LLMs.