AI Explained
Plain explanations of trending AI concepts, with live visualizations.
Sparse MoEs overfit repeated data sooner than dense models — Data-repetition tolerance — What does it mean?
MoEs start degrading at 4 data passes, dense models hold past 8, and MoEs fall behind dense after 32. Strong masking holds the lead past 64.
REVA cuts RAG compression overhead up to 15.6× — Document-keyed evidence views — What does it mean?
REVA scores a document once from past generator attention, then serves any token budget from that store — no compressor pass per query.
Compress context up to 266× with one ratio-adaptive model — Matryoshka memory budgets — What does it mean?
One soft-context compressor whose every prefix is a valid coarser summary, so each input picks its own compression ratio.
Make NVMe KV-cache loads 2x faster with scheduler-aware preloading — External KV-cache break-even — What does it mean?
An external KV-cache hit is an admission decision with a break-even, not a free lookup.
Audit finds a typical DeepSeek-V4-Flash site uses about two of its four residual streams — Near-identity late residual mixing — What does it mean?
DeepSeek-V4-Flash carries four residual streams, but across layers 22–42 it barely mixes them.
Pretrained LLMs resist 4-bit quantization partly because layer errors cancel — Counteracting quantization error — What does it mean?
Pretrained LLMs survive 4-bit rounding because each layer's new error opposes the one it inherited — a cancellation learned in pretraining.
Cut RAG TTFT 80% with selectively recomputed KV-cache chunks — Selective KV recomputation — What does it mean?
Reusing precomputed KV caches per RAG chunk is fast but drops cross-chunk attention; re-running a chosen subset buys the quality back.
Open-source Miles for asynchronous post-training across 64 GB300 GPUs — Rollout-trainer separation with asynchronous weight sync — What does it mean?
Generation and optimization are different shapes of work. Splitting them across GPU pools raises one question: how do the weights cross?
NVIDIA Dynamo — Encode-prefill-decode disaggregation — What does it mean?
Dynamo gives the vision encoder its own workers, so a text request stops queuing behind somebody else's image.
vLLM 0.29 — Batch-sharded sampling — What does it mean?
Every tensor-parallel rank used to sample the whole batch. vLLM 0.29 splits it, cutting per-step logits memory to 1/TP.
vLLM 0.29 — Mamba prefill checkpoints — What does it mean?
An SSM state cannot be sliced like a KV cache, so vLLM 0.29 snapshots it mid-prefill to make prefix caching work.
Cut Qwen3 MoE expert traffic by up to 53.3% with cache-aware routing — Predicting MoE expert demand before the layer runs — What does it mean?
A learned cache router predicts which experts a layer will need, so they are already resident when it asks.