AI Explained
Plain explanations of trending AI concepts, with live visualizations.
TypeSafe releases Jev — Typed parallel decisions — What does it mean?
Jev fixes the output type before it runs, samples every field in parallel, and attaches a calibrated probability to each one.
Emergence World runs 16-day agent worlds — Non-compositional alignment — What does it mean?
Eight 10-agent worlds, 16 days, no survivors: agents detected the injection and still acted on it up to 46 hours later.
GAUGE finds 57.5% of agent chats rated satisfying still failed the user's task — Ranking validity vs construct validity — What does it mean?
An offline judge can rank agents right and still measure the wrong thing: 57.5% of chats rated satisfying failed the user's task.
GenV catches formal encodings that pass the solver but changed the question — Verdict-preserving unfaithfulness — What does it mean?
A solver-valid encoding can still mean the wrong problem. GenV scores the gap the verdict cannot see.
The Router Within — Glance-and-verdict skill routing — What does it mean?
The signal for picking an agent skill already sits inside the frozen model's forward pass — two linear maps read it out.
Injected plans slip past chain-of-thought monitors — Plan injection — What does it mean?
A planted plan the actor restates as its own reasoning leaves the chain-of-thought monitor reading a clean, laundered trace.
A tool call can succeed while the workflow fails — Observation-effect separation — What does it mean?
A tool call's outcome can be unknown. Splitting world effects from runtime observations names the eight ways recovery then goes wrong.
Gander lets you interrupt a model mid-answer — Cerebellum-Brain split — What does it mean?
Gander runs a Cerebellum that holds the live conversation and a Brain that reasons, so you can interrupt it mid-answer.
AWS open-sources an agent stack that watches behavior and infrastructure separately — Behavioral vs infrastructural telemetry — What does it mean?
Infra telemetry proves an agent ran; behavioral telemetry proves it was right. Agents fail the second way while the first stays green.
TAM benchmark — Exact-match scoring over long procedural rule chains — What does it mean?
TAM grades the whole rule chain, not the steps: 1% exact match on ICD-10-CM coding, 15.5% on sentencing.
MCP SEP-2640 — Progressive skill disclosure — What does it mean?
MCP SEP-2640 puts each skill's frontmatter, digests and file sizes in the listing, so a host can budget before fetching a byte.
RTK reported 89% token savings and DeepSeek's cost rose 17% — Turn amplification — What does it mean?
RTK reported 89% token savings. DeepSeek's per-task cost still rose 17%, because thinner turns bought more turns.