AI Explained
Plain explanations of trending AI concepts, with live visualizations.
CheckRLM corrects factual drift inside retrieval-augmented reasoning — In-chain retrieval fact-checking — What does it mean?
CheckRLM fact-checks each claim inside a reasoning model's chain of thought against retrieved evidence, then patches only the wrong step.
Google releases TabFM for zero-shot tabular prediction — Tabular in-context learning — What does it mean?
TabFM predicts on a spreadsheet it has never seen by reading the whole table as one prompt — in-context learning, no per-dataset training.
SkillCoach self-evolves rubrics to grade agentic skill-use at scale — Self-evolving rubrics — What does it mean?
SkillCoach grades how an agent uses its skills on four axes — with a rubric that rewrites its own criteria, keeping only edits that still grade trusted samples right.
kNNGuard turns LLM hidden activations into a training-free guardrail — Training-free activation-space kNN guardrail — What does it mean?
A guardrail that flags unsafe prompts by reading a frozen LLM's hidden activations and kNN-matching a 50-prompt bank — no training, up to 10× faster.
BlockSearch traces million-token retrieval collapse — Attention dilution — What does it mean?
BlockSearch traces long-context retrieval collapse to attention dilution — the softmax denominator drowning the gold document — and fixes it with a length-aware softmax.
AgenticSTS tests long-horizon agent memory — Bounded-memory contract via typed retrieval — What does it mean?
AgenticSTS reframes agent memory as a contract — each decision sees only what typed retrieval hands it, so the prompt stays bounded and each memory layer is testable.
TRIAGE cuts agent turns up to 14.8% — Role-typed credit assignment — What does it mean?
TRIAGE grades each agent action by its role — progress, probe, stall, or blunder — not just the final win, so RL wastes fewer turns.
SkillHone evolves agent skills across sessions — Persistent decision-history memory — What does it mean?
SkillHone improves an agent's skills across sessions from an external decision history it reads — no retraining, no weight changes.
QVal: training-free testbed finds prompting beats dense agent supervision — Q-aligned dense supervision — What does it mean?
QVal tests, with no training, whether an agent's per-step supervision ranks actions like a reference policy's Q-values — and finds prompting beats 21 methods.
Orca learns a unified world model over video and events — Next-State-Prediction — What does it mean?
Orca trains one shared world model with a single objective — predict the next state — then reads it out as text, image, or action.
CausalMix picks LLM pretraining data mixtures via causal inference — CATE-based data mixture selection — What does it mean?
CausalMix picks an LLM's pretraining data mix by causal inference: estimate each mix's causal effect from 512 tiny runs, then scale to 7B.
ELDR routes MoE decode by expert locality, cutting TPOT up to 13.9% — Expert-locality-aware decode routing — What does it mean?
ELDR predicts a request's experts and routes decode to the worker whose zone matches, so fewer experts reload — cutting TPOT 5.9–13.9%.











