PoS (Progression of States) — Belief trapping vs explicit belief state — What does it mean?
The news. On October 1, 2026, researchers posted Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief States, which introduces PoS, an inference-time framework that keeps an explicit, validated belief state and recovers the agent when it stalls. Across four benchmarks — two where the agent must change the world (ALFWorld, LOCA-Bench) and two where it must diagnose a cause (RCA-100, ClinDiag) — and three LLM backbones, it reports the highest overall score in every benchmark-and-backbone pair. The claim underneath is that keeping more history is not the bottleneck; keeping an accurate, checked picture of now is. Read the paper →
Picture two detectives on the same case. The first keeps one notebook and, every morning, rereads it from page one to work out where things stand. Monday's page says the side door was shut; Wednesday's page says a witness saw it open. Nothing on any page says which one is true today — the detective has to rebuild that from the pages each time, and on a bad morning rebuilds it wrong. The second detective keeps a case board: one column for what is true now, one for what they still need to find out, one for what they still need to get done. The board is not a tidier notebook — it is a different object: an estimate of the present, not a record of the past.
The second detective also has a partner who looks at every new pin before it goes up. If someone tries to pin "door: open" next to "door: shut", the partner stops it and asks which one the evidence supports. And even with a clean board, a detective can still waste a week re-interviewing the same witness or following leads that answer none of the open questions. That second failure — moving without progress — is what PoS calls belief trapping, and a correct board alone does not prevent it. In curriculum terms, PoS turns the agent's state object from a log of what happened into a model of what is true.
Concretely, PoS keeps a belief with four parts: a structured world state (entities, their states, and relations, each with provenance and confidence), the goal, the open epistemic gaps, and the open achievement gaps. On each tick of the loop it picks one active gap to work on, so the agent does not hop between open items without finishing any. After the action, the agent writes a candidate update to the belief. The Belief Sentinel then audits the new and changed entries for two kinds of error — the update contradicts itself, or it contradicts the latest observation — and the agent must revise before the belief is committed. This targets the same rot that Context Engineering's four failure modes describe: a wrong fact that, once in context, quietly steers every later step.
Trapping detection runs on top of the committed beliefs. On execution tasks, each step gets a progress label: did it shrink the active gap, gather information needed for it, or move toward the goal? Over a rolling window of the last K = 8 validated steps, PoS computes three numbers: gap persistence, measured separately for learn-gaps and do-gaps (what fraction of the gaps of that type open at the start of the window stayed unresolved at every step of it), stagnation (what fraction of steps made no progress), and recurrence (how often the part of the world tied to the active gap returns to a state it was already in). These combine into one belief health score, H = 1 − max(P_learn, P_do) × max(S, R), and the agent is declared trapped when H ≤ 0.25. Unlike a plain retry or a stop rule, PoS then diagnoses the trap — static (nothing changes), cycle (states repeat), or drift (things change, but not the part that matters) — and whether a learn-gap or a do-gap is blocked, and adds a matching constraint to the next action: suppress the useless move, break the cycle, or re-anchor to the gap; and demand new, discriminating evidence or a real state change.
Here is how the health score separates a stuck agent from a slow one, using the paper's own window K = 8 and threshold 0.25, with illustrative gap and step counts. An agent heating a mug starts the window with two do-gaps — mug inside, mug heated — and no open learn-gaps, so the learn-gap persistence P_learn is 0 by the paper's rule. After 8 steps both do-gaps are still open, so P_do = 2 / 2 = 1.0; only 1 of the 8 steps made progress, so stagnation S = 7 / 8 = 0.875; take recurrence R = 0.5 (illustrative). Then H = 1 − max(0, 1.0) × max(0.875, 0.5) = 0.125, below 0.25 — trapped, recovery starts. Now change one fact: the mug went in during that window, so one of the two gaps closed and P_do = 1 / 2 = 0.5. With the same 7 idle steps, H = 1 − max(0, 0.5) × max(0.875, 0.5) = 0.5625 — healthy, no intervention. Health falls below the threshold only when open gaps persist and the agent is idling or revisiting old states at the same time; closing gaps pushes it back up.
| Result (Qwen3.7-Plus backbone) | Reported | Source |
|---|---|---|
| ALFWorld task success: raw trajectory → best baseline → PoS | 62.69% → 72.39% → 88.81% | Table 1, PoS paper |
| ALFWorld without the Sentinel's consistency check | 73.88% (−14.93 points) | Table 1 ablation, PoS paper |
| RCA-100 joint accuracy: raw trajectory → best baseline → PoS | 24.27% → 28.16% → 38.83% | Table 1, PoS paper |
| RCA-100 tokens per episode: raw trajectory → PoS | 355.70K → 1,800.13K (5.06×) | Table 2, PoS paper |
The bill is real. On RCA-100 with Qwen3.7-Plus, the task agent itself uses 20.9% fewer tokens than with the raw trajectory, because it reads a compact belief instead of the whole history — but building and auditing that belief brings the total to about 5× the tokens per episode. The ablations show where the money goes: dropping the Sentinel saves 35.6% of PoS's tokens and costs 7.76 accuracy points; dropping trapping diagnosis saves only 9.2% and costs 4.85 points. The paper's bet is that inference-time compute is better spent maintaining the agent's picture of the world than squeezing its history, and trapping is common enough to justify it — on RCA-100, the Kimi-K3 backbone was detected as trapped in 74% of episodes. The method was tested on benchmarks; whether the 5× overhead pays off on a given production workload is a cost question each team still has to measure.
Goes deeper in: AI Agents → The Agent Loop & State → The State Object
Related explainers
- Harness-1 — externalized state — moves a search agent's working memory out of the transcript; PoS adds checks that the externalized picture stays consistent and keeps moving
- CAPTURE — counterfactual memory auditing — audits long-term memory writes about a user; the Sentinel audits in-episode updates about the world
- Agentic abstention — when to stop — decides when to quit under uncertainty; PoS decides when the agent is stuck and how to restart progress