AI Explained

Plain explanations of trending AI concepts, with live visualizations.

GPU

Protect LLM training from silent corruption with 1.65–6.76% overhead — Silent data corruption in transformer training — What does it mean?

A silent bit flip does not crash training. TrainSDC guards only the interfaces that amplify a fault, at 1.65-6.76% runtime overhead.

LLM

Translate KV states across model families to cut prefill 67% — Cross-model KV translation — What does it mean?

A learned layer rewrites one model's KV cache into another model's format, so the next model resumes instead of prefilling: 899 to 138 ms.

LLM

vLLM 0.28 ships disk-backed tiered KV offload — Partial loads from a lower cache tier — What does it mean?

A tier below the GPU can hand back part of a cached prefix. vLLM 0.28 lets that partial answer be kept rather than thrown away.

Agent

PolicyGuide compiles agent policy into a workflow graph — Workflow-graph guidance vs local action vetoes — What does it mean?

PolicyGuide tracks where an agent is inside a policy graph, so a check returns the next compliant step instead of a refusal.

LLM

Masked diffusion serving measured at 16× batch throughput — Denoising-step batching — What does it mean?

Only 24% of a masked-diffusion request's wall clock is GPU math. Line 16 requests up on the same denoising step and throughput rises 16×.

Agent

EvoMal shows agents copy a planted payload into the skills they write — Skill-library self-poisoning by imitation — What does it mean?

Nobody runs the planted skill. The agent copies its payload while imitating its shape, files the copy, and the copy gets imitated next.

LLM

AsymSpec gives the small drafter the full context and the large verifier a compressed one — Context-asymmetric speculative decoding — What does it mean?

Speculative decoding usually shows both models the same context. AsymSpec shows the cheap model everything and the expensive model a summary.

LLM

KeysAndValues finetunes models under the KV policy they will serve with — KV policy co-adaptation — What does it mean?

Long-context models train seeing every past token, then get served seeing a fraction of them. This paper applies the eviction policy during finetuning instead.

GPU

NVIDIA pairs Rubin GPUs with Groq 3 LPX — Phase-split inference across two accelerator types — What does it mean?

Prefill is compute-bound and decode is memory-bandwidth-bound. NVIDIA now runs each phase on a different kind of chip inside one rack.

GPU

Microsoft details the Maia 200 inference accelerator — Software-defined dataflow — What does it mean?

A GPU decides at runtime what runs next. Maia 200 moves that into the compiler, which writes the data-movement schedule ahead of time.

LLM

Show chunked prefill beats elastic KV-cache reclamation — Reclaiming the prefill reserve — What does it mean?

A serving engine holds GPU memory back for the next prefill. A clever allocator can lend it out — a smaller prefill chunk frees more.

GPU

Target 10,000 tok/s with Cerebras CS-5 and 3D-stacked memory in CS-6 — Wafer-scale memory locality — What does it mean?

Decode re-reads the weights for every token. Cerebras' answer is to stop moving them: keep the weights on the same wafer as the math.