OnTrack interrupts failing agent runs — Streaming trajectory alignment — What does it mean?
The news. On October 8, 2026, Babak Barazandeh, Connor Swanson, Chinmay Kulkarni and Nikhil Mungel of the Cribl AI Research Lab posted OnTrack: Real-Time Monitoring and Intervention in LLM Agent Trajectories via Streaming Structure-Aware Optimal Transport. It takes an earlier after-the-fact scoring method, Otap, and makes it run on a live event stream. On held-out runs of the SWE-agent coding agent on SWE-bench (real GitHub bug-fix tasks), it ranks failing runs below successful ones better than content similarity after the first 8 steps (+0.057 AUROC), and an abort policy built on top of it saves about 18% of the compute spent on failing runs. Read the paper →
Picture a GPS that knows the routes other drivers took to reach the same address. It does not wait until you arrive to grade the trip, and it does not phone a second driver at every junction for an opinion. It compares the road you have driven so far with the roads that worked, and it speaks up when you start circling a block. OnTrack is that GPS for an agent: it compares the run so far with recorded successful runs while the run is still going. The run is modeled as an execution graph: each step (an action, a tool call, its arguments and its result) is a node, and an edge joins two steps when a later step uses something an earlier step produced. If the agent searches at step 5 and edits the file it found at step 15, the edge (5, 15) appears. Reference runs are stored as graphs of the same kind. It is the trace you already get from one span per agent tick, with the dependencies drawn in.
To compare the two graphs, OnTrack uses optimal transport. Each agent step has a small amount of mass to spread over reference steps, and sending mass to a reference step costs more the more the two steps differ in what they say and which tool they call. A second term compares structure: two steps that sit close together in the agent's graph should land on steps that sit close together in the reference. A step whose mass cannot be placed cheaply anywhere "leaks": it can signal a digression, a step taken too early or a hallucinated action, though a leak alone does not say which. The alignment itself comes from Otap, which scored finished runs. The paper's contribution is running it on a stream, and that is harder than re-running the batch scorer after every step, for four reasons.
| Streaming problem | What goes wrong if you re-run the batch scorer | OnTrack's fix | Source |
|---|---|---|---|
| Prefix vs whole | The agent is penalized for every reference step it has not reached yet — for being early, not wrong | Frontier masking: expect only the part of the reference the agent could have reached so far | §4.2 |
| Provisional structure | A new step looks disconnected until a later step uses its output, so it looks like a hallucination | A new step starts with zero structural weight and a grace window of 3 events; late edges are inserted in place | §4.4, §4.5 |
| Compute budget | Solving the full alignment again at every step grows with the run and misses the deadline of the next tool call | One warm-started solver iteration per event, a median of 0.85 ms per event at 40 steps and 3 references on one Apple M5 Max core; a converged solve only before an intervention and once at the end of the run | §4.3 |
| Decisions, not scores | One bad step is 1/5 of a 5-step run but a small dent at step 30, so no fixed score threshold means the same thing at every length | Decisions come from per-step signals; the overall score is reported but never thresholded | §1, §4.5 |
The GPS also keeps working when it knows less. With references and tool schemas (each tool's declared inputs and whether it is reversible), OnTrack can flag a deviation from the plan and block an irreversible action whose prerequisites are missing; with schemas only, it still catches loops, stalls and repeated calls and keeps that gate; with nothing but the event stream, loop and stall detection remain. The signals are simple per-step checks: a loop score for a near-duplicate of an earlier action (retries of a failed step are exempt up to a limit), coverage velocity for activity without progress, and information gain for new calls that produce no new state. The pre-execution gate is a policy check on the proposed tool call, so it runs before the call executes and can block a deletion whose declared prerequisites are missing, instead of reporting it afterwards; what it can catch depends on how well the schemas are annotated.
A flag is not the same as a stop. The naive rule, abort at the first flag, is useless: nearly every run, successful or not, raises an early exploration flag, so it stops 99% of runs. The paper counts only severe flags (loop, off-track, a step taken before its prerequisites exist, a blocked call, a stall) and aborts when their density in a rolling window of 6 steps passes a threshold. Here is how the headline numbers compose. Hold fixed the corpus mix, where 81% of runs fail, and the illustrated threshold of 0.60, which stops 23.3% of failing runs and 20.7% of successful runs on a balanced sample. Out of 1,000 runs, 810 fail and 190 succeed, so the policy stops about 0.233 × 810 ≈ 189 failing runs and 0.207 × 190 ≈ 39 successful runs. Precision is 189 / (189 + 39) ≈ 0.83: about five of every six aborts hit a run that was going to fail, and the early stops save 17.9% of the compute the failing runs would have used. The authors say 0.60 was chosen for exposition after the sweep, not on held-out data, that the false-stop rate ranged from 13% to 21% across two samples, and that a deployment should set its own threshold on its own traffic.
The advantage is an early one. OnTrack beats embedding similarity at ranking runs only in the first ~10 steps; by step 15 plain cosine similarity catches up (0.633 vs 0.635 AUROC). That fits the job, because a monitor that saves compute must decide early, and the cost profile of an agent is driven by runs that keep calling the model. It also relates to compounding errors in multi-step agents, although the paper does not measure how early mistakes spread. Two limits matter. OnTrack detects process anomalies that correlate with failure; it does not predict success, and on SWE-agent's narrow search–read–edit–test template the structure term added little over a step-by-step attribute match. And a plain step-count cutoff saved comparable cost at a matched ~20% false-stop rate; the authors argue OnTrack's edge is a stable threshold and a reason attached to each stop ("loop at step t"), which allows a softer response such as replanning or a human escalation.
Goes deeper in: Agent Engineering → Observability for Agents → Alerting on Agent Behavior
Related explainers
- Agentic CLEAR grades agents at three zoom levels — another approach to grading an agent after the run
- White-box deception probes — a different live monitor that reads hidden states instead of the trace
- Idempotency keys for exactly-once tool effects — what to do about the irreversible actions a pre-execution gate is meant to protect