AI Explained

Plain explanations of trending AI concepts, with live visualizations.

Agent

TypeSafe releases Jev — Typed parallel decisions — What does it mean?

Jev fixes the output type before it runs, samples every field in parallel, and attaches a calibrated probability to each one.

Agent

Emergence World runs 16-day agent worlds — Non-compositional alignment — What does it mean?

Eight 10-agent worlds, 16 days, no survivors: agents detected the injection and still acted on it up to 46 hours later.

Agent

GAUGE finds 57.5% of agent chats rated satisfying still failed the user's task — Ranking validity vs construct validity — What does it mean?

An offline judge can rank agents right and still measure the wrong thing: 57.5% of chats rated satisfying failed the user's task.

Agent

GenV catches formal encodings that pass the solver but changed the question — Verdict-preserving unfaithfulness — What does it mean?

A solver-valid encoding can still mean the wrong problem. GenV scores the gap the verdict cannot see.

Agent

The Router Within — Glance-and-verdict skill routing — What does it mean?

The signal for picking an agent skill already sits inside the frozen model's forward pass — two linear maps read it out.

Agent

Injected plans slip past chain-of-thought monitors — Plan injection — What does it mean?

A planted plan the actor restates as its own reasoning leaves the chain-of-thought monitor reading a clean, laundered trace.

Agent

A tool call can succeed while the workflow fails — Observation-effect separation — What does it mean?

A tool call's outcome can be unknown. Splitting world effects from runtime observations names the eight ways recovery then goes wrong.

Agent

Gander lets you interrupt a model mid-answer — Cerebellum-Brain split — What does it mean?

Gander runs a Cerebellum that holds the live conversation and a Brain that reasons, so you can interrupt it mid-answer.

Agent

AWS open-sources an agent stack that watches behavior and infrastructure separately — Behavioral vs infrastructural telemetry — What does it mean?

Infra telemetry proves an agent ran; behavioral telemetry proves it was right. Agents fail the second way while the first stays green.

Agent

TAM benchmark — Exact-match scoring over long procedural rule chains — What does it mean?

TAM grades the whole rule chain, not the steps: 1% exact match on ICD-10-CM coding, 15.5% on sentencing.

Agent

MCP SEP-2640 — Progressive skill disclosure — What does it mean?

MCP SEP-2640 puts each skill's frontmatter, digests and file sizes in the listing, so a host can budget before fetching a byte.

Agent

RTK reported 89% token savings and DeepSeek's cost rose 17% — Turn amplification — What does it mean?

RTK reported 89% token savings. DeepSeek's per-task cost still rose 17%, because thinner turns bought more turns.