AI Explained

Plain explanations of trending AI concepts, with live visualizations.

Agent

Where Does Exactly-Once Live? — Idempotency keys vs read-back verification for exactly-once tool effects — What does it mean?

A timed-out agent write may still land. Read-back checks miss it; idempotency keys cut duplicates from 28% to 4% in the LIMBO benchmark.

Agent

Taste-Bench: best model picks the right fork 59.7% of the time — Outcome-hidden decision forks — What does it mean?

Taste-Bench scores an agent's choices at decision forks before the outcome shows; the best of 14 frontier models gets 59.7%.

Agent

Critical-State RL picks which model call to train — Nested continuation sampling — What does it mean?

Before training a tool agent, measure which call's action actually moves the reward — then train only that call.

Agent

CliffCompaction — Re-compacting from the original context — What does it mean?

Compact only the turns since the last compaction, keep text verbatim, discard the old compaction: no summary-of-summary drift.

Agent

NemotronLabs VoiceChat listens while it speaks — Full-duplex vs half-duplex turn-taking — What does it mean?

A full-duplex speech model listens and speaks on parallel channels, so silence becomes a trained output rather than a gap between turns.

Agent

ActObs unmasks the environment's half of an agent trajectory — Observation-token supervision — What does it mean?

Agent SFT grades only the agent's commands. ActObs also supervises the environment's replies - about 45% of the trajectory tokens.

Agent

Agent skill memory that rewrites itself from user traffic — Matched replay gate — What does it mean?

An agent rewrites its own skill library, and a matched replay gate refuses any edit that regresses a single replayed case.

Agent

A 176-setting study ablates the coding-agent harness — Harness component ablation — What does it mean?

Benchmarking a whole harness hides which part earned the gain. Ablate one component and the answer changes with model and budget.

Agent

LangChain benchmarks Jev — Evaluator variance vs the generative LLM judge — What does it mean?

A grader has error of its own. Freeze the agent's trace, replay the grader, and the spread you see is entirely the grader's.

Agent

Schedule agent tools by task-specific CPU bottlenecks — Task-aware CPU admission vs uniform resource allocation — What does it mean?

An agent's bottleneck is often local CPU or disk, not the model — so size each task's container by what its tools actually need.

Agent

Chronicle replays agent incidents as CI tests — Cut-point replay vs stubbing every boundary — What does it mean?

Freeze the boundaries you are not testing, run the one you are live, and an incident becomes a CI test.

Agent

Magenta closes informal math reasoning with Lean-guided repair — Error-attribution routing — What does it mean?

Magenta names why a Lean proof failed, then routes the retry to a re-derivation or a one-step patch instead of resampling.