AI Explained
Plain explanations of trending AI concepts, with live visualizations.
Allocate low-bit recurrent state precision by channel decay — Decay-aware bit allocation for recurrent states — What does it mean?
A rounding error in a recurrent state lasts as long as its row remembers, so slowly decaying rows get more bits at the same average.
CacheReforge — Stale KV cache repair after LoRA adapter updates — What does it mean?
A LoRA update makes cached K/V stale. CacheReforge recomputes about 5% of layers instead of half and removes 92% of the output drift.
Reparameterize output heads before W4 quantization — Softmax-preserving head shift — What does it mean?
Softmax ignores shared logit shifts, so choosing an equivalent output head before rounding can make a 4-bit head far more faithful.
Crossflow — Revocable decode-node leases — What does it mean?
Let decode GPUs lend short-lived, revocable capacity to prefill when prompt demand spikes, instead of re-splitting a fixed P/D cluster.
From Memory Budgets to Risk Targets — Risk-controlled KV cache eviction — What does it mean?
Pick how much KV cache to evict by certifying how often requests break, not by average score, and keep the full cache if nothing passes.
Greedy decoding is not precision-invariant — Top-two logit margin decides BF16 vs FP16 flips — What does it mean?
Same model, same prompt, greedy decoding: BF16 and FP16 still disagree, and a tiny top-two logit margin decides where.
CompKV sparse attention paper — Compensation-aware KV block selection — What does it mean?
CompKV picks which KV blocks sparse attention reads exactly by the error a block-mean stand-in would leave: mass times logit variance.
Tree speculative decoding on DeepSeek-V4 — Branch-isolated compressed state — What does it mean?
Tree speculation on compressed attention needs a temporary state per branch and a commit of only the accepted path.
KV-COBRA splits the KV-cache budget per head — Rank vs bit-width allocation — What does it mean?
KV-COBRA picks rank and bit width per attention head instead of globally, so the same KV-cache bits land where they cut the most error.
The Undetected Damage of Quantization on Retrieval — Top-1 score-gap certificate — What does it mean?
A quantized model's top answer is certified only if it leads the runner-up by at least 2x the score shift; retrieval rarely does.
Complex KDA extends Kimi Delta Attention — Signed gates for state tracking — What does it mean?
Negative gates plus a full-strength delta rule make two mirrors, and two mirrors make a rotation: one linear-attention update that can count in cycles.
vLLM 0.30 ships Fast Start — Persistent GPU weight cache via CUDA IPC — What does it mean?
vLLM 0.30 keeps prepared weights in a per-GPU daemon, so a restarted engine maps them over CUDA IPC instead of reloading them from disk.