Agent·

Causal Memory Policy paper — Randomized memory exposure with inverse-propensity weighting — What does it mean?

The news. On October 1, 2026, Arman Behnam and Binghui Wang of the Illinois Institute of Technology posted Causal Memory Policy: Making Memory Utility Identifiable by Intervening on Retrieval (arXiv 2610.02070). It argues that memory systems which decide what to keep by estimating each memory's effect on answers are blind to every memory their retriever never surfaces, and it measures how large that blind spot is: 54% of the memories a question needs were never retrieved on LongMemEval, and 67% on LoCoMo — two benchmarks that test recall across long multi-session conversations. Its fix reserves a few context slots for memories drawn on a fixed, known schedule. Read the paper →

Picture the bookstore. Each month the manager decides which titles to reorder by reading the sales ledger. The front table holds six books, and it always shows the current best sellers. A book that has never been on the front table has never had a chance to sell, so its sales column reads zero. The ledger cannot tell a bad book from a book that nobody ever saw — and a manager who reads zero sends that book back to the publisher.

A memory-augmented agent runs the same shop. Its retriever ranks the stored memories and puts the top few into the context window. The retention policy then decides which memories to keep by asking how much each one improved the answers. For a memory that was never retrieved, that question has no data behind it, and the usual estimators quietly return a number close to zero.

The root cause is that retrieval sits between the memory store and the answer. A memory can change an answer only if it is in the context the model reads. Recent systems estimate usefulness by changing the store — keep a memory or remove it, then compare answers. That is store-level randomization, and if the retriever would not have picked the memory anyway, both versions of the store produce the same context and the same answer. The contrast is zero, which looks exactly like "this memory is useless." In causal-inference terms this is a positivity violation, and it happens at the retrieval step, so checks that only look at memory operations never see it.

The paper's LongMemEval numbers show how one-sided retrieval support is. Of 113 required memories, 61 are never retrieved, 6 more almost never are, 37 are retrieved on more than 80% of draws, and only 9 sit in between. Store-level randomization left the retrieved set unchanged on 52.2% of interventions. A larger retrieval window does not fix it: in the deployed Mem0 system, surfacing half of the never-retrieved required memories took a window of 138 memories, and 90% took 286 — at that size the system is no longer retrieving, it is pasting most of the store. This is a RAG failure mode that sits in how the system is evaluated, not in a single answer.

CMP fixes the experiment design, not the estimator: it randomizes retrieval itself. Out of B = 6 context slots, k = 2 are taken away from the ranker. The other 4 still hold the top-ranked memories, and the 2 reserved slots are filled from an exposure pool (a candidate list of memories eligible for the rotation) by a balanced schedule — the bookstore giving 2 of its 6 front-table spots to a rotation. Every pool memory gets the same number of exposures (give or take one), and those exposures are assigned to turns at random. Because the counts are fixed in advance, each memory's chance of being shown, its propensity, is known exactly instead of guessed from logs.

The usefulness estimate is a self-normalized inverse-propensity-weighted (Hájek) contrast: average answer quality on the turns where the memory was shown, minus average quality on the turns where it was not. With a balanced schedule the weights are constants, so the estimator reduces to a plain difference of two means with an exact variance. It is the same idea as an A/B harness: you learn what a change does only when you assign it on purpose. One gap remains: a memory the ranked slots retrieve on every turn never has an absent arm, so the paper adds a forced-exclusion variant for that case.

DesignWhat is randomizedNever-retrieved memory gets both arms?Required vs non-required AUC (LongMemEval)
Logged retrieval (observational)Nothing; propensities are estimated from logsNo — positivity failsNot identified for never-retrieved memories (paper)
Store-level randomizationWhich memories are in the storeNo — the ranker still skips it0.542 (paper, Table 2)
CMP exposure (k = 2 of B = 6)Which memories fill 2 context slotsYes — by construction0.664 (paper, Table 2)

Here is the arithmetic of the schedule, with the paper's LongMemEval settings held fixed: B = 6 slots, k = 2 reserved, T = 90 turns, a pool of about 32 memories (take exactly 32). That gives 90 × 2 = 180 exposure slots to share among 32 memories, or 5.6 each. A balanced schedule gives 20 memories 6 exposures and 12 memories 5 exposures (20 × 6 + 12 × 5 = 180), so every propensity is known: 6/90 ≈ 0.067 or 5/90 ≈ 0.056. Drawing the slots independently with the same expected count would leave roughly a third of the pool below five exposures, under the paper's five-per-arm bar. Now take one memory that the ranked top 4 never retrieve, shown in a rotation slot on 6 turns and absent on 84. Suppose (illustrative) its answers average a token F1 of 0.50 on the 6 shown turns and 0.20 on the other 84: its estimated utility is 0.50 − 0.20 = +0.30. Under store-level randomization with a retriever that never surfaces it, the with-memory and without-memory runs read the same context, and the estimate is 0 — the number that gets a memory evicted.

On LongMemEval, the redesign raises AUC from 0.542 to 0.664 and halves the damage that eviction does. Gold loss — required memories destroyed per 100 context slots reclaimed — falls from 10.9 to 5.2, and required memories are nominated for eviction at 0.68 times the chance rate instead of 0.97. When the authors split memories by retrieval support, the whole gain sits in the never-retrieved group; where both designs observe both arms, they score about the same. The cost is small but real: reserving 2 of 6 slots lowered mean token F1 from 0.175 to 0.165 across 51 items, a difference that was not statistically significant (p = 0.62) and came mostly from two items where the ranker was already retrieving the right evidence. That is the scarce-context trade: every slot spent on measurement is a slot not spent on the best guess.

The paper is explicit that measuring a memory's utility is not the same as knowing whether to keep it. Per-query utility reaches 0.78 AUC on the question it was measured for, yet no aggregation across queries that a retention policy could use predicts a memory's value on unseen questions. The experiments also use an oracle exposure pool — each question's required memories plus the 20 highest-ranked and 10 random others — and the three practical pool-selection rules the authors tested had recall near zero. So deciding which memories deserve a rotation slot is still an open problem. Read CMP as a measurement design for memory evaluation, the error-analysis step of a memory system, rather than a ready eviction policy.

Goes deeper in: AI Agents → Retrieval & RAG → RAG Failure Modes

Related explainers

Frequently Asked Questions

Check what you knowMap your AI & GPU knowledge across every track — free, role-based