RECAST — Adaptive evidence routing between retrieval and computation — What does it mean?
The news. On October 7, 2026, researchers from MIT and Google posted RECAST (Routing Evidence through Computation, Access, and Synthesized Tools, arXiv 2610.10507). They trained a Qwen3.5-9B model to build evidence for a frozen Gemini 3.5 Flash answer model, and tested it on six benchmark families whose sources range from dataframes and financial reports to Wikipedia passages and user profiles. The trained system reached 75.6% average success, 15.9 percentage points above the strongest baseline, and it improved over the strongest baseline by 15.0 points on average on three benchmarks it never trained on. Read the paper →
Picture an executive who asks a research assistant: which three-month period had the largest increase in operating margin? (Operating margin is operating income divided by revenue.) The filing cabinet holds monthly revenue and monthly operating income. No folder in the cabinet contains the answer, because the answer is a number that must be calculated from many folders. An assistant who can only pull folders will bring back a stack of monthly reports and leave the executive to do the division and the comparisons. This is the paper's own motivating example, and it is the gap in a plain retrieve-then-generate pipeline: retrieval finds text that exists, it does not create text that does not.
A good assistant has three moves. For a known name or ID, they search the cabinet by keyword or by meaning. For a filter or a total, they run a spreadsheet formula over the records. For something unusual, they ask the IT desk for a one-off script, and they must describe the job completely, because the IT desk never sees the executive's question. After each move they look at what came back, and they hand over the brief only when it is enough.
RECAST turns those moves into an action menu for the RouterLM. First, a rule-based preprocessor (no language model) converts the source into uniform records and writes a short source profile: how many records there are, their format and their fields. The RouterLM sees that profile, never the whole source. Then, in each round, it picks one action:
- CALL_PRIMITIVE: a Lexical (BM25) search, a Semantic (embedding) search, or a Relational read-only SQL query, with the query and the number of records to return. A call can be limited to records found in an earlier round.
- SYNTHESIZE: a written specification that the frozen CompilerLM turns into one Python program. The program can read the complete source, reuse earlier evidence and call the primitives.
- ACCEPT_CONTEXT: stop, and pass the collected evidence to the frozen AnswerLM.
After every call the executor reports whether the operation ran, how many records matched and any error message. A call that runs without errors is not the same as evidence that is sufficient, so the decision to stop is a separate judgment that the RouterLM must learn. In the experiments the loop is capped at six rounds and 4,000 characters of kept evidence. This is the chaining-and-routing pattern applied to evidence collection, with when to stop as a learned decision.
Combining search and SQL with synthesized code gives a higher average score than allowing either family alone. With the untrained Qwen router, allowing only primitives gave 56.8% average success and allowing only synthesized code gave 59.9%; allowing both gave 64.1%. The trained policy also learned a different mix for each source. On DataBench (questions over dataframes) it ordered a custom program on 98% of questions; on HotpotQA (Wikipedia passages), no question used synthesis in the one run the paper reports; over all six families it used synthesis on 27.3% of questions and took 3.32 rounds on average. So the learned policy treats code as a tool for some sources, not a default: on passage collections it kept to retrieval, which still leaves it exposed to the classic RAG failure modes.
Training changes who does the expensive thinking. The RouterLM is first fine-tuned (SFT) on 3,368 training rows built from its best successful runs, with rarer benchmark families repeated to balance the mix. It is then trained with GRPO on 608 rows drawn from 567 unique questions that it solved only sometimes. GRPO keeps only groups of four attempts that contain at least one success and one failure, so every group has a contrast to learn from. The reward is 0.90 × answer correctness + 0.08 × token-level F1 + 0.02 × valid action structure, where token-level F1 measures word overlap with the reference answer. The small weights matter: when two attempts in a group are both wrong, they still let GRPO tell the nearly right, well-formed one from the other. Removing that filter dropped the average from 75.6% to 73.0%, and using correctness alone as the reward dropped it to 68.9%.
| Approach | What builds the evidence | Avg. success | Tokens per question |
|---|---|---|---|
| Fixed retrieval | One embedding search, top 3 chunks | 55.7% (Table 1) | 1.6k (Table 13) |
| Direct code (Gemini 3.5 Flash) | One Python program over the source, no second round | 45.1% (Table 1) | 4.8k (Table 13) |
| IRCoT (Gemini 3.5 Flash) | Reasoning steps alternating with retrieval | 58.8% (Table 1) | 10.0k (Table 13) |
| Interact-RAG (Gemini 3.5 Flash) | Planner, reasoner and executor that plan retrieval and decide when to stop | 59.7% (Table 1) | 26.9k (Table 13) |
| RECAST, trained Qwen3.5-9B router | Retrieval, SQL or a synthesized program, chosen each round | 75.6% (Table 1) | 19.4k (Table 13) |
Where the tokens go. Token counts here include input and output across every model call except the evaluation judge. Hold three things fixed, as the paper does: the same Gemini 3.5 Flash AnswerLM for every method, the same 100 test questions per benchmark family, and the same frozen judge. The best baseline, Interact-RAG with a Gemini controller, uses 26.9k tokens per question, all of them on the large model, and succeeds 59.7% of the time. Trained RECAST uses 13.9k tokens on the small Qwen router plus 5.5k on Gemini (for the CompilerLM and the AnswerLM), a total of 19.4k. So 26.9k − 19.4k = 7.5k fewer tokens, which is 27.9% fewer in total, and 26.9k − 5.5k = 21.4k fewer large-model tokens, which is 79.5% fewer, while success rises by 15.9 points to 75.6%. The saving comes from moving the many small routing decisions to a cheap trained model and keeping the large model for the two jobs that need it: writing code and writing the answer.
Two limits are worth keeping in mind. All scores come from an LLM judge (Gemini 3.5 Flash), though the paper reports that two other judges and token-level F1 keep RECAST ahead. And the hardest family stays hard: on MultiHiertt (several hierarchical tables plus text) the trained system reaches 65.3%, the best in that column, while the same router without training reached only 29.3%. The gains vary by benchmark, and the paper itself notes that repeated router calls and occasional code synthesis can add latency and cost compared with a single retrieval pass. Building the training data also costs extra compute, because every candidate trajectory has to be executed and judged.
Goes deeper in: AI Agents → Retrieval & RAG → Retrieve-Then-Generate
Related explainers
- Attention dilution in long-context retrieval — the other way to skip a retriever: paste everything into the context, and why that breaks at scale
- Contextual-bandit tool routing — learning which tool to call, the per-call cousin of RECAST's per-round evidence routing
- Code-as-action interface — agents that act by writing code, where RECAST writes code only to build evidence