AI Explained
Plain explanations of trending AI concepts, with live visualizations.
Route local agent inference across mixed devices with NVIDIA PAIR — Request-level routing — What does it mean?
PAIR routes each whole inference request to one local device instead of splitting it, so more agent jobs run at once, not faster.
Merge per-task GRPO experts to serve 116M monthly requests — Two-stage SLERP merging of per-axis experts — What does it mean?
Train one GRPO expert per capability axis, then merge them along the sphere in two stages instead of one multi-objective run.
IFM releases K2 Horizon — Uno's LoRA diffusion adapter for block-parallel decoding — What does it mean?
Uno freezes K2 Horizon's weights and clips on a LoRA diffusion adapter that writes whole token blocks per pass.
Let language models declare which KV-cache regions to attend — Model-declared attention scope vs external sparsity predictors — What does it mean?
The model writes global, focus or local inside its own reasoning; the engine parses it like a tool call and skips most of the KV cache.
Spend quantization bits globally instead of repairing critical layers — Global quantization granularity — What does it mean?
A 9-model causal study finds quantization damage is diffuse: at a small matched budget, a finer group size beats repairing a few layers.
Translate KV states across model families to cut prefill 67% — Cross-model KV translation — What does it mean?
A learned layer rewrites one model's KV cache into another model's format, so the next model resumes instead of prefilling: 899 to 138 ms.
vLLM 0.28 ships disk-backed tiered KV offload — Partial loads from a lower cache tier — What does it mean?
A tier below the GPU can hand back part of a cached prefix. vLLM 0.28 lets that partial answer be kept rather than thrown away.
Masked diffusion serving measured at 16× batch throughput — Denoising-step batching — What does it mean?
Only 24% of a masked-diffusion request's wall clock is GPU math. Line 16 requests up on the same denoising step and throughput rises 16×.
AsymSpec gives the small drafter the full context and the large verifier a compressed one — Context-asymmetric speculative decoding — What does it mean?
Speculative decoding usually shows both models the same context. AsymSpec shows the cheap model everything and the expensive model a summary.
KeysAndValues finetunes models under the KV policy they will serve with — KV policy co-adaptation — What does it mean?
Long-context models train seeing every past token, then get served seeing a fraction of them. This paper applies the eviction policy during finetuning instead.
Show chunked prefill beats elastic KV-cache reclamation — Reclaiming the prefill reserve — What does it mean?
A serving engine holds GPU memory back for the next prefill. A clever allocator can lend it out — a smaller prefill chunk frees more.
Relation reports lower loss than MHA at 10M–100M params — Self and Exchange relations — What does it mean?
Attention normalizes pairwise scores into weights immediately. Relation splits that same evidence into Self and Exchange channels first.





