Agent·

Where Does Exactly-Once Live? — Idempotency keys vs read-back verification for exactly-once tool effects — What does it mean?

The news. On September 24, 2026, Jiapeng Li posted Where Does Exactly-Once Live? Model, Harness, and Tool-Contract Effects on Duplicate Side Effects in LLM Agents to arXiv. The paper introduces LIMBO, a deterministic sandbox of six services with twelve fault modes injected at the service boundary, and grades every episode against a ledger of the effects that were actually committed. It runs 25,930 episodes across nine models, three production agent harnesses and two tool-contract variants to ask where exactly-once behaviour should be enforced: in the model, in the harness, or in the tool contract. Read the paper →

Picture paying a bill by mailing a cheque and then hearing nothing. You phone the bank, and it has no record of the payment. Should you mail another cheque? If the first cheque was lost, a second one is correct; if it is still in the mail, a second one pays the bill twice, and from your chair the two cases look the same. That is the position of an agent whose write call returned a timeout or a 500 error. The request may never have reached the service, may have succeeded with its reply lost, or may still be queued and about to land. A retry policy that only asks "did it fail?" cannot choose correctly, because the failure signal is the same in all three cases.

The paper separates the cases by what a phone call to the bank can reveal. When the write already landed and only the reply was lost (a lost acknowledgement), a read-back sees the effect, and the right move is to stop. When the request is still in flight (a late commit) or the transport delivers it twice (redelivery), the read-back finds nothing, or finds one copy while a second one is on its way. For those faults a more careful agent does not help: the paper proves that no recovery policy that relies only on what it can observe is exactly-once under late commits, unless it knows an upper bound on how long a request can stay in flight.

The fix that works lives in the tool, not in the agent: write the cheque number on the cheque. An idempotency key is a unique ID the client sends with a write; the service records it and answers any repeat of the same key with the original result instead of doing the work again. If the retry reuses the first attempt's key, the late original becomes a no-op in every case, so the agent no longer needs to know which case it is in. This is the same argument the idempotency step makes for recovering a crashed harness, applied at the boundary between an agent and a remote service.

The measured results follow the proof. Frontier models told to act exactly once almost never duplicated a write whose acknowledgement was lost (0.5% of episodes), because the read-back showed the truth; for that kind of fault the model was the main factor, explaining 53% of the explained variance. When the request was still in flight or was redelivered, the same frontier models duplicated in 56% and 74% of episodes, and the tool contract explained 81% of the variance. Which layer owns exactly-once depends on the fault: the model can own the faults a read-back can see, and only the tool contract can own the rest.

Two details make the gap worse in practice. A read path that cannot see lagging or in-flight writes was worse than having no read path at all, because agents trusted its empty answer instead of escalating to a human. And transparent retries added by client-side middleware, which the model never sees, cut exactly-once success from 72% to 50%, because a hidden retry of a non-idempotent write turns every lost acknowledgement into a duplicate.

FaultWhat a read-back showsFrontier duplicate rateWho can fix itSource
Lost acknowledgementThe write, already done0.5%The model: check before retryingAbstract
Late commit (in flight)Nothing yet56%The contract: a key, or a known in-flight boundAbstract
RedeliveryCannot stop the second copy74%The contract: a keyAbstract

Walk the paper's waiting experiment, where in-flight delays are heavy-tailed (most requests land quickly, a few take a very long time). The alternative to keys is to wait long enough for any in-flight request to land, then read back. Waiting 5 minutes reached 67% exactly-once success; waiting 1 hour reached 84%, at 49.9 minutes of agent-visible time per episode, because a one-hour bound covered the delay in only 82% of the simulated cases. The keys-everywhere contract, with a guard that attaches keys, reached 94% in 1.5 minutes. Holding the delay distribution fixed, keys gave 10 more points of exactly-once success in about 1/33 of the time (49.9 ÷ 1.5 ≈ 33, our arithmetic). Even an oracle that knows exactly when each in-flight request will land reached 99% only at 19.5 minutes. The paper's contract test shows the same effect in counts: when every write accepted a key, agents attached one in 98% of episodes and the duplicate rate fell from 28% to 4%, about 24 fewer duplicates per 100 episodes in that comparison. Every remaining duplicate came from an agent that sent the first attempt without a key or changed the key when it retried, so a key protects only when it is created once per intended action and reused on every retry.

Two more findings matter for anyone running agents in production. First, agents do not know when they have duplicated an effect: in 90% of the episodes that produced a duplicate, the agent reported the task as completed, and in 80% it listed no operation as uncertain, so a person reading the final report had no reason to check. That makes duplicates a what-to-log problem: the ledger of committed effects is the evidence, not the agent's own summary. Second, the harness barely mattered. Three production harnesses and a minimal scaffold running the same model behaved almost identically, and a guard that attaches keys worked across harnesses unchanged.

The limits are the benchmark's own: six simulated services with injected faults rather than production traffic, and per-model numbers tied to the models tested in September 2026. The durable recommendation is aimed at tool and protocol designers: every non-idempotent tool should accept a standard idempotency-key argument, declare a read-back operation, and document how long its writes can stay invisible or in flight. That is now a question worth asking in any tool design review.

Goes deeper in: Agent Engineering → Production Harness Architecture → Idempotency — Safe to Re-run

Related explainers

Frequently Asked Questions

Check what you knowMap your AI & GPU knowledge across every track — free, role-based