AI Explained
Plain explanations of trending AI concepts, with live visualizations.
GAUGE finds 57.5% of agent chats rated satisfying still failed the user's task — Ranking validity vs construct validity — What does it mean?
An offline judge can rank agents right and still measure the wrong thing: 57.5% of chats rated satisfying failed the user's task.
GenV catches formal encodings that pass the solver but changed the question — Verdict-preserving unfaithfulness — What does it mean?
A solver-valid encoding can still mean the wrong problem. GenV scores the gap the verdict cannot see.
NVLink 6 multi-layer resiliency — Recovery escalation ladder — What does it mean?
NVLink 6 fixes a fault on the cheapest rung that can: FEC under 1 ms, software 1.5 s, shadow engine 7.3 s, cold restart 283 s.
FlexEE cuts LLM decoding by up to 3.16x under weight offloading — KV-compatible early exit — What does it mean?
Stopping a token at layer 12 saves compute and a weight fetch — but only if the layers it skipped still leave a usable KV cache.
OPEN-1B publishes a training run you can replay bit for bit — Step-replay auditing — What does it mean?
OPEN-1B orders every reduction, batch and collective so one training step — not the whole run — becomes the unit anyone can verify.
Cartridges match in-context learning once retrieval is real — KV cartridges vs parametric fine-tuning — What does it mean?
In multi-document retrieval, a trained KV cartridge is the only method that matches leaving the documents in the context window.
LoopSpec drafts the next token while it verifies this one — Pipelined self-speculative decoding — What does it mean?
A looped Transformer can read a draft off its own early passes, and LoopSpec makes that draft happen while the full depth verifies.
The Router Within — Glance-and-verdict skill routing — What does it mean?
The signal for picking an agent skill already sits inside the frozen model's forward pass — two linear maps read it out.
Injected plans slip past chain-of-thought monitors — Plan injection — What does it mean?
A planted plan the actor restates as its own reasoning leaves the chain-of-thought monitor reading a clean, laundered trace.
A tool call can succeed while the workflow fails — Observation-effect separation — What does it mean?
A tool call's outcome can be unknown. Splitting world effects from runtime observations names the eight ways recovery then goes wrong.
MAPS cuts LLM tail latency up to 84.8% — Uncertainty-calibrated output-length bounds — What does it mean?
Output length is unknown when a request arrives. MAPS hands the scheduler a ceiling with a chosen error rate instead.
Gander lets you interrupt a model mid-answer — Cerebellum-Brain split — What does it mean?
Gander runs a Cerebellum that holds the live conversation and a Brain that reasons, so you can interrupt it mid-answer.