TokenCast — Segment-level token forecasting — What does it mean?
The news. On September 28, 2026, researchers from Sun Yat-Sen University, Rensselaer Polytechnic Institute, Southeast University and HKUST posted TokenCast (arXiv 2609.35760). They collected 11,712 execution traces from 240 tasks across SWE-bench Verified, Search-R1, MMLU-Pro and LongBench-v2, run by six agent models, and trained forecasters that update after every call. The paper reports a 14.5% average cut in forecast error against the strongest comparator over 96 benchmark-model-checkpoint combinations, and 21.3% fewer tokens than a fixed budget in an offline stopping replay. Read the paper →
Picture the binder again. Before every meeting, the office copies the whole binder for the room, so the copy bill for a meeting is the binder's thickness on that day. A page added early is not paid for once; it is paid for again at every meeting that follows. That is the shape of an agent's bill. Each LLM call re-sends the full conversation, so a 2,000-token tool output that lands in call 3 shows up again in the input of call 4, call 5, and every call after it. In the paper's words, “the growing context steadily inflates the input size of every subsequent call.”
This is why a total is hard to guess up front. The number of meetings is not known in advance: a test fails, the agent retries, a file is bigger than expected. Each of those choices changes both how many calls are left and how thick the binder will be for all of them, which is the double effect that makes context a scarce resource in the first place.
TokenCast's move is to describe cost per segment, not per run. A segment A is a block of calls that starts with input length LA, spans nA calls and costs CA tokens. The paper records it as three numbers: the call count nA, the context growth gA, and a residual bA = CA − nA·LA, which is everything beyond re-reading the starting binder (generated output plus growth inside the segment).
When segment B follows segment A, the paper shows the combined residual is exactly bA + bB + nB·gA, and that last term is the bill for re-reading A's pages during every call of B. Nothing is estimated in that step; it is bookkeeping. The estimating happens elsewhere: LightGBM models predict the next segment's numbers and the state it will end in, a second model predicts the rest of the run from that boundary, and the identity stitches the two together. A separate direct model predicts the remaining total in one shot, and a small correction model blends both paths.
The four points above are where TokenCast refreshes. The forecast improves as evidence arrives, and the composition path only wins once the run has started. The paper scores this as normalized MAE: the forecast's error divided by the error of a naive predictor that always guesses the historical median, so 1.0 means no better than that guess and lower is better. Before the first call, a direct model is more accurate than composition (0.82 vs 0.87). After calls have completed, the order flips (0.66 vs 0.72), and averaging the two gives 0.64. Removing context growth from the segment description raises the full method's normalized MAE from 0.69 to 0.76, so the re-read term is doing measurable work. Every refresh is a set of tree-model lookups, which the paper measures at a mean of 32.8 ms per run, summed over all refreshes, on SWE-bench Verified, with no extra LLM call. That matters for observability: the forecast can sit next to each call span without adding a cost of its own.
Here is the re-read term with numbers (illustrative). An agent starts a task with a 20,000-token context: system prompt, task and a repository map. Segment A is its first 2 calls, during which it reads one file and the output adds 8,000 tokens to the context (assume nothing else is added), so gA = 8,000. Segment B is the next 10 calls, so nB = 10. The cross term is nB·gA = 10 × 8,000 = 80,000 input tokens. In words, a single 8,000-token file read costs about 80,000 tokens by the end of the run, ten times its own size, and a forecaster that only looked at call sizes would miss all of it. If the agent then retries and adds 5 more calls, that one file costs another 40,000.
The payoff is in deciding when to stop. A fixed budget is the paper cap: the run keeps going until it has spent the cap, and only then is it stopped, even if it was never going to finish. TokenCast's controller instead adds confirmed spend to a chosen percentile of the forecast remaining spend (the 5th, 50th or 95th, picked on held-out runs) after each call, and stops the run as soon as that sum passes the budget. In a replay of 288 recorded GPT-5.4 runs on SWE-bench Verified, that rule used 21.3% fewer tokens on average across seven budgets while matching the fixed budget's trace-completion rate at every budget. At the 174k-token budget, for example, both reached 30.2% trace completion, at 154.7k tokens per run for the fixed cap and 101.1k for TokenCast. The replay checks whether a run reaches its recorded end state, not whether the task was solved, so read these as token savings at equal completion, not as a success-rate result.
One thing the forecast does not change is the price of a token. Prompt caching lets a provider bill a re-read prefix at a discount on the major APIs, but the tokens are still metered. TokenCast predicts provider-accounted token counts, so the two are complementary: caching makes each re-read cheaper, and the forecast tells you how many re-reads are still coming.
| Approach | What it predicts | Counts re-read context? | Extra work per forecast |
|---|---|---|---|
| Fixed token budget | Nothing; stops at the cap | No | None |
| Output-length predictors (TRAIL, EGTP, TIE) | Length of one model reply | No, one reply at a time | A local encoder or model-state pass (paper App. D.7) |
| Self-Prediction | Total or remaining task tokens, estimated by the agent model itself | Only as the model judges it | Extra calls to the agent LLM (paper App. D.7) |
| TokenCast | This call, and the rest of the run, refreshed after each call | Yes, the nB·gA term | Tree-model lookups, 32.8 ms per run (paper) |
The limits are worth stating. The forecasters are trained on traces from specific harnesses (DeepSeek Harness and OpenHands), and on independently released LiveClawBench trajectories from a domain they had not seen, they did worse than Self-Prediction (MAE ratios of 1.31 at Call Start and 1.47 at Task Update), recovering to 0.82 and 0.85 after 20 target-domain tasks. TokenCast also trailed the strongest comparator in 24 of the 96 combinations, 15 at Task Start and 9 at Call Start, and in none at the two checkpoints that come after output has been observed. The durable lesson is independent of the model: an agent's cost is set less by how long each reply is than by how long each early addition stays in the context.
Goes deeper in: Agent Engineering → Cost & Latency Engineering → The Cost Profile of an Agent
Related explainers
- Token Budgets — affine-typed budget ownership — enforces a spending cap at compile time instead of forecasting spend at run time
- EarlyEval — calibrated early stopping — predicts a run's verdict mid-way; TokenCast predicts its remaining cost
- Turn amplification behind RTK — a case where thinner turns meant more turns, so the per-call view misled
- Cost per successful outcome — the metric to divide a forecast by once you know the success rate