AI Explained

Plain explanations of trending AI concepts, with live visualizations.

Agent

ContextRL rewards evidence selection to boost agent and multimodal reasoning — Contrastive context-selection RL — What does it mean?

ContextRL rewards a model for picking which of two near-identical contexts supports the answer — sharpening fine-grained evidence grounding.

GPU

UFP4 fixes FP4 pretraining's shrinkage bias — E2M1 shrinkage bias — What does it mean?

E2M1's lopsided 4-bit bins round values toward zero — a shrinkage bias UFP4 fixes with a Hadamard transform + stochastic rounding.

LLM

Taylor-Calibrate cuts hybrid-attention distillation tokens 4.9–9.2× — Taylor-guided gate initialization — What does it mean?

Taylor-Calibrate presets a linear-attention student's gates from the teacher — hitting distillation targets with 4.9–9.2× fewer tokens.

Agent

FAPO auto-optimizes multi-step LLM pipelines, beating GEPA on 15 of 18 benchmarks — Failure-attribution-gated prompt optimization — What does it mean?

FAPO has Claude Code diagnose where an LLM pipeline fails, then make scoped prompt or chain edits — beating GEPA on 15 of 18 benchmarks.

Agent

AtomMem gives LLM agents memory built from atomic facts, SOTA on LoCoMo — Atomic-fact agent memory — What does it mean?

AtomMem distills an agent's long history into atomic facts, files them by event and time, and links them in an associative graph for retrieval.

Agent

LedgerAgent gives tool-calling agents a structured state ledger — Pre-tool-call policy validation — What does it mean?

LedgerAgent tracks an agent's task state in a separate ledger and checks domain policy against it before any irreversible tool call.

LLM

HydraHead fuses full and linear attention per head, not per layer — Head-axis attention hybridization — What does it mean?

HydraHead mixes full and linear attention head-by-head, keeping exact attention only for retrieval-critical heads — a 7:1 split matching a coarser 3:1.

LLM

EfficientRollout — Self-speculative decoding with quantized self-drafters — What does it mean?

EfficientRollout speeds RL rollouts by drafting with a quantized copy of the model itself — a self-drafter that tracks the evolving policy for free.

LLM

CacheWeaver reorders RAG evidence for prefix-cache reuse — Prefix-cache-aware evidence reordering — What does it mean?

CacheWeaver reorders the retrieved chunks in a RAG prompt so the serving engine reuses its KV prefix cache — cutting median TTFT 20–33% with no measured quality loss.

Agent

Agent leaderboards mislead under distribution shift (IBM) — Predictive validity — What does it mean?

IBM: agent leaderboards rank models by one aggregate score that fails under distribution shift — measure predictive validity instead.

LLM

LoopCoder-v2: two loops of a shared block beat deeper looping — Weight-tied block looping — What does it mean?

Running one shared transformer block twice lifts a 7B model from 43.0 to 64.4 on SWE-bench — but three or more loops regress, a non-monotonic sweet spot.

LLM

ConSA learns where to put full vs sliding-window attention per head — Controllable attention sparsity — What does it mean?

ConSA learns, under a fixed budget, which attention heads get full attention and which get a cheap sliding window — instead of hand-coded rules.