AI Explained
Plain explanations of trending AI concepts, with live visualizations.
Taste-Bench: best model picks the right fork 59.7% of the time — Outcome-hidden decision forks — What does it mean?
Taste-Bench scores an agent's choices at decision forks before the outcome shows; the best of 14 frontier models gets 59.7%.
Critical-State RL picks which model call to train — Nested continuation sampling — What does it mean?
Before training a tool agent, measure which call's action actually moves the reward — then train only that call.
Greedy decoding is not precision-invariant — Top-two logit margin decides BF16 vs FP16 flips — What does it mean?
Same model, same prompt, greedy decoding: BF16 and FP16 still disagree, and a tiny top-two logit margin decides where.
CompKV sparse attention paper — Compensation-aware KV block selection — What does it mean?
CompKV picks which KV blocks sparse attention reads exactly by the error a block-mean stand-in would leave: mass times logit variance.
CliffCompaction — Re-compacting from the original context — What does it mean?
Compact only the turns since the last compaction, keep text verbatim, discard the old compaction: no summary-of-summary drift.
Tree speculative decoding on DeepSeek-V4 — Branch-isolated compressed state — What does it mean?
Tree speculation on compressed attention needs a temporary state per branch and a commit of only the accepted path.
KV-COBRA splits the KV-cache budget per head — Rank vs bit-width allocation — What does it mean?
KV-COBRA picks rank and bit width per attention head instead of globally, so the same KV-cache bits land where they cut the most error.
The Undetected Damage of Quantization on Retrieval — Top-1 score-gap certificate — What does it mean?
A quantized model's top answer is certified only if it leads the runner-up by at least 2x the score shift; retrieval rarely does.
Complex KDA extends Kimi Delta Attention — Signed gates for state tracking — What does it mean?
Negative gates plus a full-strength delta rule make two mirrors, and two mirrors make a rotation: one linear-attention update that can count in cycles.
vLLM 0.30 ships Fast Start — Persistent GPU weight cache via CUDA IPC — What does it mean?
vLLM 0.30 keeps prepared weights in a per-GPU daemon, so a restarted engine maps them over CUDA IPC instead of reloading them from disk.
AMDKernelVault trains an 8B model to write AMD GPU kernels — Hierarchical execution reward — What does it mean?
Binary pass/fail leaves most GRPO rounds with zero advantage. AMDKernelVault grades a kernel 0.4, 1.0 and up to 1.5 instead.
Qwen-Image-2.1 runs two mask rules in one sequence — Mixed-granularity attention — What does it mean?
One sequence, two mask granularities: causal for text, chunk-level for image — plus a static-context KV cache computed once.