AI Explained

Plain explanations of trending AI concepts, with live visualizations.

Agent

Taste-Bench: best model picks the right fork 59.7% of the time — Outcome-hidden decision forks — What does it mean?

Taste-Bench scores an agent's choices at decision forks before the outcome shows; the best of 14 frontier models gets 59.7%.

Agent

Critical-State RL picks which model call to train — Nested continuation sampling — What does it mean?

Before training a tool agent, measure which call's action actually moves the reward — then train only that call.

LLM

Greedy decoding is not precision-invariant — Top-two logit margin decides BF16 vs FP16 flips — What does it mean?

Same model, same prompt, greedy decoding: BF16 and FP16 still disagree, and a tiny top-two logit margin decides where.

LLM

CompKV sparse attention paper — Compensation-aware KV block selection — What does it mean?

CompKV picks which KV blocks sparse attention reads exactly by the error a block-mean stand-in would leave: mass times logit variance.

Agent

CliffCompaction — Re-compacting from the original context — What does it mean?

Compact only the turns since the last compaction, keep text verbatim, discard the old compaction: no summary-of-summary drift.

LLM

Tree speculative decoding on DeepSeek-V4 — Branch-isolated compressed state — What does it mean?

Tree speculation on compressed attention needs a temporary state per branch and a commit of only the accepted path.

LLM

KV-COBRA splits the KV-cache budget per head — Rank vs bit-width allocation — What does it mean?

KV-COBRA picks rank and bit width per attention head instead of globally, so the same KV-cache bits land where they cut the most error.

LLM

The Undetected Damage of Quantization on Retrieval — Top-1 score-gap certificate — What does it mean?

A quantized model's top answer is certified only if it leads the runner-up by at least 2x the score shift; retrieval rarely does.

LLM

Complex KDA extends Kimi Delta Attention — Signed gates for state tracking — What does it mean?

Negative gates plus a full-strength delta rule make two mirrors, and two mirrors make a rotation: one linear-attention update that can count in cycles.

LLM

vLLM 0.30 ships Fast Start — Persistent GPU weight cache via CUDA IPC — What does it mean?

vLLM 0.30 keeps prepared weights in a per-GPU daemon, so a restarted engine maps them over CUDA IPC instead of reloading them from disk.

GPU

AMDKernelVault trains an 8B model to write AMD GPU kernels — Hierarchical execution reward — What does it mean?

Binary pass/fail leaves most GRPO rounds with zero advantage. AMDKernelVault grades a kernel 0.4, 1.0 and up to 1.5 instead.

LLM

Qwen-Image-2.1 runs two mask rules in one sequence — Mixed-granularity attention — What does it mean?

One sequence, two mask granularities: causal for text, chunk-level for image — plus a static-context KV cache computed once.