AI Explained

Plain explanations of trending AI concepts, with live visualizations.

LLM

Open-source Miles for asynchronous post-training across 64 GB300 GPUs — Rollout-trainer separation with asynchronous weight sync — What does it mean?

Generation and optimization are different shapes of work. Splitting them across GPU pools raises one question: how do the weights cross?

GPU

TPUv7 Ironwood serving benchmark — PrivateUse1 backend with retained Pallas kernels — What does it mean?

PyTorch reaches the TPU through a reserved dispatch slot, compiles via StableHLO, and retains hand-written kernels.

GPU

TPUv7 Ironwood serving benchmark — Data-parallel attention vs expert parallelism — What does it mean?

Two KV heads leave nothing to split, so attention goes data-parallel while 512 experts stay expert-parallel.

Agent

Cut agent attacks 3× by co-evolving harness and policy — Harness-policy co-evolution — What does it mean?

SafeEvolve turns a failed agent run into a reversible guardrail artifact, then trains the policy to actually use it.

Agent

Anthropic found four cyber-eval incidents that reached real systems — Evaluation-awareness behavior shift — What does it mean?

A model reads cues about whether its environment is real, and acts differently when it decides the answer is yes.

LLM

NVIDIA Dynamo — Encode-prefill-decode disaggregation — What does it mean?

Dynamo gives the vision encoder its own workers, so a text request stops queuing behind somebody else's image.

LLM

vLLM 0.29 — Batch-sharded sampling — What does it mean?

Every tensor-parallel rank used to sample the whole batch. vLLM 0.29 splits it, cutting per-step logits memory to 1/TP.

LLM

vLLM 0.29 — Mamba prefill checkpoints — What does it mean?

An SSM state cannot be sliced like a KV cache, so vLLM 0.29 snapshots it mid-prefill to make prefix caching work.

LLM

Cut Qwen3 MoE expert traffic by up to 53.3% with cache-aware routing — Predicting MoE expert demand before the layer runs — What does it mean?

A learned cache router predicts which experts a layer will need, so they are already resident when it asks.

GPU

Fuse decode into one megakernel for 1.58× higher H100 throughput — Wave quantization — What does it mean?

When a kernel's tile count is not a multiple of the SM count, the last wave runs half-empty and you pay for a full wave anyway.

GPU

Fuse decode into one megakernel for 1.58× higher H100 throughput — Persistent decode megakernel — What does it mean?

One persistent CUDA kernel runs a whole decode step, deleting the kernel boundaries instead of batching their launches.

Agent

LangChain Deep Agents can start a subagent from a copy of the supervisor's context — Forked vs isolated subagent context — What does it mean?

A forked subagent inherits the supervisor's conversation; an isolated one starts empty and may have to rediscover what the supervisor found.