AI Explained
Plain explanations of trending AI concepts, with live visualizations.
ContextRL rewards evidence selection to boost agent and multimodal reasoning — Contrastive context-selection RL — What does it mean?
ContextRL rewards a model for picking which of two near-identical contexts supports the answer — sharpening fine-grained evidence grounding.
UFP4 fixes FP4 pretraining's shrinkage bias — E2M1 shrinkage bias — What does it mean?
E2M1's lopsided 4-bit bins round values toward zero — a shrinkage bias UFP4 fixes with a Hadamard transform + stochastic rounding.
Taylor-Calibrate cuts hybrid-attention distillation tokens 4.9–9.2× — Taylor-guided gate initialization — What does it mean?
Taylor-Calibrate presets a linear-attention student's gates from the teacher — hitting distillation targets with 4.9–9.2× fewer tokens.
FAPO auto-optimizes multi-step LLM pipelines, beating GEPA on 15 of 18 benchmarks — Failure-attribution-gated prompt optimization — What does it mean?
FAPO has Claude Code diagnose where an LLM pipeline fails, then make scoped prompt or chain edits — beating GEPA on 15 of 18 benchmarks.
AtomMem gives LLM agents memory built from atomic facts, SOTA on LoCoMo — Atomic-fact agent memory — What does it mean?
AtomMem distills an agent's long history into atomic facts, files them by event and time, and links them in an associative graph for retrieval.
LedgerAgent gives tool-calling agents a structured state ledger — Pre-tool-call policy validation — What does it mean?
LedgerAgent tracks an agent's task state in a separate ledger and checks domain policy against it before any irreversible tool call.
HydraHead fuses full and linear attention per head, not per layer — Head-axis attention hybridization — What does it mean?
HydraHead mixes full and linear attention head-by-head, keeping exact attention only for retrieval-critical heads — a 7:1 split matching a coarser 3:1.
EfficientRollout — Self-speculative decoding with quantized self-drafters — What does it mean?
EfficientRollout speeds RL rollouts by drafting with a quantized copy of the model itself — a self-drafter that tracks the evolving policy for free.
CacheWeaver reorders RAG evidence for prefix-cache reuse — Prefix-cache-aware evidence reordering — What does it mean?
CacheWeaver reorders the retrieved chunks in a RAG prompt so the serving engine reuses its KV prefix cache — cutting median TTFT 20–33% with no measured quality loss.
Agent leaderboards mislead under distribution shift (IBM) — Predictive validity — What does it mean?
IBM: agent leaderboards rank models by one aggregate score that fails under distribution shift — measure predictive validity instead.
LoopCoder-v2: two loops of a shared block beat deeper looping — Weight-tied block looping — What does it mean?
Running one shared transformer block twice lifts a 7B model from 43.0 to 64.4 on SWE-bench — but three or more loops regress, a non-monotonic sweet spot.
ConSA learns where to put full vs sliding-window attention per head — Controllable attention sparsity — What does it mean?
ConSA learns, under a fixed budget, which attention heads get full attention and which get a cheap sliding window — instead of hand-coded rules.











