Planning-as-Routing — The plan declaration–execution gap — What does it mean?
The news. On September 29, 2026, researchers at ADIA Lab posted Do LLM Agents Execute the Plans They Declare? (arXiv 2609.38108). They test three models (Qwen3.6-35B-A3B, DeepSeek-V4-Flash and Gemma-4-26B) on ALFWorld, Mind2Web, WebArena and SWE-bench Verified, under three conditions that share one model and one action space: Flat ReAct with no plan, Plan+ReAct, and Planning-as-Routing. A rule-based verifier checks whether each Plan+ReAct trajectory actually reaches the declared plan steps in order. Read the paper →
Picture the taxi. You hand the driver a four-stop itinerary, and the driver reads it once and then drives the way taxi drivers drive: look at the street, pick the next turn, repeat. Nothing in a taxi forces stop 2 to come before stop 3, so whether the itinerary survives depends on the driver keeping it in mind at every corner. That is a ReAct loop with a plan pasted into its prompt. Each tick, the model chooses its next action from the latest observation, and the plan is one more piece of context competing for attention.
The paper measures how often the stops actually happen. The verifier strips each declared plan to its actionable steps, matches them to the actions in the trajectory, and counts a run as structure maintained only if every scorable step appears in the declared order (extra actions in between are allowed). These are verifier-based estimates: on 100 ALFWorld trajectories its agreement with two human annotators was κ = 0.31 and 0.34, about the same as the humans' agreement with each other (κ = 0.35), so the matching is imperfect. On ALFWorld, Mind2Web and SWE-bench Verified, the authors summarise the rate as only about 22–45%, and one model on SWE-bench Verified fell to about 12%. The gap widens with plan length: on ALFWorld, plans of 4–5 steps kept their structure in roughly 37–50% of runs, 6–7-step plans in roughly 5–21%, and longer plans almost never. Short plans tend to survive the taxi and long ones mostly do not (the paper flags some long-plan bins as sparse, and it measures the association, not the cause). It looks like the compounding-error problem applied to a plan instead of to individual steps.
The travel desk works differently. The model still chooses the kind of trip, but a deterministic router then puts the task on an executor whose code is that kind of trip, with no extra model call to decide the dispatch. A Predefined executor writes the whole plan and runs it without replanning. A Sequential executor runs one step, observes, and revises the remaining plan. A Hierarchical executor is the orchestrator–workers pattern: it splits the task into subgoals and hands each to a worker. A Search executor generates several candidate plans, runs each one, and has a rubric-based judge pick the best trajectory, close to the parallelization and voting pattern. In the paper's words, this makes "that structure part of the execution control flow". Structurally it is routing: one classification up front, then a specialised handler.
Splitting the run into a declaration and an execution is what makes failures attributable, and that is the second half of the idea. With Plan+ReAct, a failed task could mean the model chose a bad plan or followed a good plan badly, and the success rate cannot tell you which. The paper separates them with forced dispatch: run every mode on every task, record a task-by-mode success matrix, and compare three numbers. The best single mode applied to everything shows what the executors can do. The model's own choice (Routing@1) shows what selection achieves. The per-task oracle shows the ceiling. This is error analysis built into the architecture rather than done by reading traces afterwards.
| Planning mode | What the executor's control flow enforces | ALFWorld success when forced on all 134 tasks (DeepSeek-V4) |
|---|---|---|
| Predefined | Full plan written first, then run with no replanning | 0.556 (Table 3) |
| Sequential | Execute a step, observe, revise the rest of the plan | 0.633 (Table 3) |
| Hierarchical | Orchestrator splits into subgoals; workers execute; results are aggregated | 0.840 (Table 3) |
| Search | Several candidate plans each executed; a rubric-based judge keeps the best trajectory | 0.918 (Table 3) |
| Plan+ReAct (baseline) | Plan in the prompt; generic ReAct loop; nothing enforced | 0.480 (Table 5) |
Here is how the three numbers compose on one cell of the results: ALFWorld, DeepSeek-V4-Flash, 134 tasks, means over three seeds, so the task counts are rounded. Plan+ReAct solves 0.480 of the tasks, about 64 of 134. Letting the model declare its own mode and routing it to the matching executor (Routing@1) solves 0.721, about 97. Forcing the Search executor on every task solves 0.918, about 123. So routing is worth about 32 more solved tasks on average than the plan-in-the-prompt run, and the best fixed mode is worth about 26 more than routing: the model often declares a mode that is not the strongest one for this benchmark. (These are differences between mean success rates, not lists of specific tasks.) Table 5 shows the same shape across the paper's grid: on each benchmark's main success metric, Routing@1 is below the best fixed mode in 11 of the 12 benchmark–model cells. Routing closes most of the execution gap; choosing the right mode is still an open problem.
Three caveats keep the headline honest. First, the widely quoted 0.48 → 0.92 compares Plan+ReAct with the best fixed mode (Search run on every task), not with the model's own routing, which reached 0.721. Second, Search executes several candidate plans and uses a judge, so it spends more rollouts. The authors control for this by running Flat ReAct three times with the same judge, and Search still led by roughly 0.12–0.39 on ALFWorld (the main text and the appendix give slightly different ranges); in their cost appendix they also report that token count and LLM-call count do not track success one-to-one. Third, the per-task oracle looks like large headroom for smarter selection, but on ALFWorld, Mind2Web and SWE-bench Verified, after three retries of the strongest single mode the remaining gap was only about −0.025 to +0.027, which suggests much of that headroom comes from extra attempts, not task-specific fit. The gains are also uneven: SWE-bench Verified moved from 0.360 to 0.442 (Hierarchical) for DeepSeek-V4, and on Mind2Web plain Flat ReAct had the highest task success of the non-oracle methods for DeepSeek-V4. The paper also finds that a judge's plan-quality score barely predicts success on ALFWorld (task-level correlation of at most 0.181), a reminder of the single-shot failure: a plan that a judge rates highly can still fail the task, so the plan text alone is a weak signal.
Goes deeper in: AI Agents → Workflow Patterns → Chaining + Routing
Related explainers
- LoopArena — controller–worker separation — what changes when the planner and the worker are separate agents, not separate executors
- SIMMER — latent failures in LLM planning — plan steps that run without error yet still undermine the goal
- Magenta — error-attribution routing — the same move of routing a failure by its cause, applied to retries