AI Explained
Plain explanations of trending AI concepts, with live visualizations.
Protect LLM training from silent corruption with 1.65–6.76% overhead — Silent data corruption in transformer training — What does it mean?
A silent bit flip does not crash training. TrainSDC guards only the interfaces that amplify a fault, at 1.65-6.76% runtime overhead.
Translate KV states across model families to cut prefill 67% — Cross-model KV translation — What does it mean?
A learned layer rewrites one model's KV cache into another model's format, so the next model resumes instead of prefilling: 899 to 138 ms.
vLLM 0.28 ships disk-backed tiered KV offload — Partial loads from a lower cache tier — What does it mean?
A tier below the GPU can hand back part of a cached prefix. vLLM 0.28 lets that partial answer be kept rather than thrown away.
PolicyGuide compiles agent policy into a workflow graph — Workflow-graph guidance vs local action vetoes — What does it mean?
PolicyGuide tracks where an agent is inside a policy graph, so a check returns the next compliant step instead of a refusal.
Masked diffusion serving measured at 16× batch throughput — Denoising-step batching — What does it mean?
Only 24% of a masked-diffusion request's wall clock is GPU math. Line 16 requests up on the same denoising step and throughput rises 16×.
EvoMal shows agents copy a planted payload into the skills they write — Skill-library self-poisoning by imitation — What does it mean?
Nobody runs the planted skill. The agent copies its payload while imitating its shape, files the copy, and the copy gets imitated next.
AsymSpec gives the small drafter the full context and the large verifier a compressed one — Context-asymmetric speculative decoding — What does it mean?
Speculative decoding usually shows both models the same context. AsymSpec shows the cheap model everything and the expensive model a summary.
KeysAndValues finetunes models under the KV policy they will serve with — KV policy co-adaptation — What does it mean?
Long-context models train seeing every past token, then get served seeing a fraction of them. This paper applies the eviction policy during finetuning instead.
NVIDIA pairs Rubin GPUs with Groq 3 LPX — Phase-split inference across two accelerator types — What does it mean?
Prefill is compute-bound and decode is memory-bandwidth-bound. NVIDIA now runs each phase on a different kind of chip inside one rack.
Microsoft details the Maia 200 inference accelerator — Software-defined dataflow — What does it mean?
A GPU decides at runtime what runs next. Maia 200 moves that into the compiler, which writes the data-movement schedule ahead of time.
Show chunked prefill beats elastic KV-cache reclamation — Reclaiming the prefill reserve — What does it mean?
A serving engine holds GPU memory back for the next prefill. A clever allocator can lend it out — a smaller prefill chunk frees more.
Target 10,000 tok/s with Cerebras CS-5 and 3D-stacked memory in CS-6 — Wafer-scale memory locality — What does it mean?
Decode re-reads the weights for every token. Cerebras' answer is to stop moving them: keep the weights on the same wafer as the math.









