AI Explained
Plain explanations of trending AI concepts, with live visualizations.
ReCache reuses tool-schema KV blocks across agent calls — Composition-invariant KV blocks — What does it mean?
Prefix caching only pays if the prompt matches token for token. ReCache encodes each tool schema on its own, so any order still hits.
MemOps benchmarks agent memory as lifecycle operations — Memory lifecycle operations — What does it mean?
MemOps grades an agent's long-term memory as four operations — remember, forget, update, reflect — not one static store of facts.
Agent optimizer study shows regression control compounds gains — Regression control in continual optimization — What does it mean?
An agent optimizer keeps compounding gains only when regression control steers it past shortcuts that boost new tasks by erasing old wins.
MCPEvol-Bench measures agents as MCP servers change — Tool-interface drift benchmarking — What does it mean?
MCPEvol-Bench mutates real MCP servers so tool interfaces drift underneath an agent — even the best of 12 models still loses about 14% once the tools evolve.
Harness Handbook localizes agent behavior before edits — Behavior-centric harness map — What does it mean?
One behavior lives in the prompt, the state, the tool call, and the coordination code. File search returns files. A behavior map returns all 4.
Plan evaluator study exposes omission incentives — Deletion non-monotonicity — What does it mean?
Deleting a step a plan needs can raise its score — the scorer drops both the step's cost and its chance of failing.
Long-Horizon-Terminal-Bench grades agent progress densely — Partial-reward threshold — What does it mean?
Pass/fail turns a 231-episode run into one bit. A partial-reward threshold grades how far the agent got, and moving the line moves the score.
OpenAI trains GPT-Red to harden agents against prompt injection — Self-play red-teaming — What does it mean?
OpenAI trained an attacker model to invent prompt injections, then fed everything it found back into the defender's training — a red-teamer that never ships.
E3 cuts coding-agent scope before expanding context — Estimate-Execute-Expand — What does it mean?
E3 is a coding-agent strategy: start from the smallest scope, expand only when a check fails, cutting cost ~85% at the same success.
Interaction scaling grounds agent feedback loops — Instrument-grounded feedback loops — What does it mean?
An agent keeps improving when it acts, measures the result with a tool, and revises — not by thinking longer or trusting a review by eye.
GATS plans agent tasks with zero LLM calls during the search — World-model tree search — What does it mean?
GATS plans an agent's actions by searching a learned world model — the search makes zero LLM calls, and it solved every task in the paper's test.
Danus coordinates math agents with fact-graph memory — Verifier-gated fact graph — What does it mean?
Danus lets parallel math agents build long proofs by gating every claim through a verifier into one shared fact graph.











