Dynamic LLM Routers are Often Misguided — Cost-matched random routing baseline — What does it mean?
The news. In early October 2026, researchers posted Dynamic LLM Routers are Often Misguided (arXiv 2610.02762). They sent 800 evaluation queries from 17 benchmarks in eight categories through six commercial routers in 14 settings — OpenRouter, Azure Model Router, Not Diamond, vLLM Semantic Router, Nadir and OrcaRouter — and generated and graded each chosen model's answer the same way. A weighted coin choosing between Gemini 3.7 Flash and Claude Opus 5 (high) at matched cost significantly beat every router except vLLM Semantic Router, which the authors had fit on their own training data; some routers trailed it by more than 10 percentage points. Read the paper →
Picture a restaurant with a host at the pass and two people at the stove: a line cook who is fast and cheap, and a head chef who is slow and expensive. The host reads each order ticket and sends it to the cheapest cook who will get it right. That is what a dynamic LLM router is: a classifier in front of a roster of models that picks one model per query, so that spend tracks how hard each query is — the routing pattern applied to cost.
Now give the host a rival: a weighted coin. The coin never reads a ticket. It sends, say, 30% of tickets to the chef and the rest to the line cook, and you tune that percentage until the coin's payroll equals the host's. If the host cannot cook more correct dishes than the coin on the same payroll, reading the tickets added nothing. That coin is the cost-matched random baseline. Its accuracy and its cost are both the weighted average of the two models' numbers, so sliding the percentage traces a straight line between the cheap model and the strong one — and any router worth paying for has to sit above that line.
Why the host loses to the coin
The paper finds four recurring failure patterns, and three of them are exactly what the standard scorecard pays for. That scorecard is Pareto efficiency on realized cost: more accuracy for less money, with each query charged what its answer actually cost.
- Difficulty blindness. The chef beats the cook most on moderately hard dishes. On the easiest, both succeed; on the hardest, both often fail — so escalating the hardest queries buys little accuracy per dollar. In the paper's data, escalating the 60th–80th percentile band of difficulty outperformed escalating the hardest 20%, and no router's choices correlated more than weakly with query difficulty.
- Length reversal. Every correct answer counts once for accuracy, but a long answer from the expensive model costs far more. So the scorecard rewards sending the chef the quick dishes, which tend to be the easy ones. Five commercial router settings showed a significant reversal of this kind.
- Semantic matching. Many routers send nearly a whole data source — all the maths, all the coding — to one model. Source is easier to predict than difficulty, and under realized-cost scoring it scores as well or better, so a router has little reason to learn difficulty at all.
- Roster suboptimality. Most routers spread traffic across many models that are not on the cost-accuracy frontier. Even a router given the true difficulty of every query, choosing its roster on the test set itself, gained at most 1 pp from rosters larger than two models.
Where the headroom actually goes
Hold three numbers fixed. From the paper: Gemini 3.7 Flash trails Claude Opus 5 (high) by 1.6 percentage points of accuracy at a 96% discount. For round numbers, say Opus costs $100 per 1,000 queries (illustrative), so Flash costs $4. Set the coin to 50/50. Its cost is 0.5 × $100 + 0.5 × $4 = $52 per 1,000 queries, and its accuracy is the average of the two models: 0.8 pp below Opus. So at roughly half the strong model's price, a router that ignores every query already lands within 0.8 pp of the strong model. If the strong model's accuracy is the ceiling, a perfect router at that budget has only about 0.8 pp left to win — and the authors' own difficulty-aware router, built on this exact pair, stayed within noise of the coin at every budget they tested. That is why the paper concludes that the roster, not the router, decides nearly all of the accuracy — the first number to read in an agent's cost profile.
The length trap is just as concrete. In one of the paper's escalation experiments, sending the 20% of queries with the shortest answers to the expensive model reached 59.5% accuracy at $1.03 per 1,000 queries — only 1.6 pp below escalating the hardest 20%, at 1.5% of that price. The authors call this accuracy hacking: calling the smarter model when it is cheapest to call, not when it is needed. Optimizing accuracy against realized cost can encourage this behavior, and the scorecard rewards it.
| Pattern | What the router does | Why realized-cost scoring rewards it | The paper's fix |
|---|---|---|---|
| Difficulty blindness | rarely escalates the hardest queries | both models often fail them, so escalating buys little accuracy per dollar | optionally reward a high hard-band escalation rate, assuming users prefer a stronger model to attempt even unsolvable queries |
| Length reversal | escalates short-answer queries | short answers are cheap to escalate and count the same for accuracy | charge each answer its model's mean cost, not its realized cost |
| Semantic matching | sends a whole data source to one model | source is easy to predict and tracks answer length | score pass-rate predictions within each data source |
| Roster suboptimality | spreads traffic over many off-frontier models | not rewarded directly; it adds ways to route wrongly | in the tested model sets, two or three well-chosen models closely approximated the best accuracy of difficulty-based routing |
How to test your own router
The portable lesson is to change the test before you trust the router. Put the router in an A/B harness against a weighted coin over your two best-value models at the same spend; if it does not clearly win, the roster is the more likely culprit, so prune it before tuning the classifier. When you compare, charge each model its mean price per answer and check escalation on your genuinely hard queries separately. The study evaluates single-turn, general-purpose benchmark queries, which may not resemble real user traffic; multi-turn routing was outside its scope, and future models may specialise more, giving routing more to do.
Goes deeper in: Agent Engineering → Cost & Latency Engineering → The Cost Profile of an Agent
Continue in trackProduction Evals — the A/B harness that runs the coin testRelated explainers
- Cluster, Route, Escalate — cost-aware LLM cascade — a router-plus-escalation design that this paper's coin test is built to check
- Co-failure ceiling — routing, voting, and mixture-of-agents — why no router can recover the queries every model gets wrong