AI Explained

Plain explanations of trending AI concepts, with live visualizations.

LLM

Route local agent inference across mixed devices with NVIDIA PAIR — Request-level routing — What does it mean?

PAIR routes each whole inference request to one local device instead of splitting it, so more agent jobs run at once, not faster.

LLM

Merge per-task GRPO experts to serve 116M monthly requests — Two-stage SLERP merging of per-axis experts — What does it mean?

Train one GRPO expert per capability axis, then merge them along the sphere in two stages instead of one multi-objective run.

LLM

IFM releases K2 Horizon — Uno's LoRA diffusion adapter for block-parallel decoding — What does it mean?

Uno freezes K2 Horizon's weights and clips on a LoRA diffusion adapter that writes whole token blocks per pass.

LLM

Let language models declare which KV-cache regions to attend — Model-declared attention scope vs external sparsity predictors — What does it mean?

The model writes global, focus or local inside its own reasoning; the engine parses it like a tool call and skips most of the KV cache.

LLM

Spend quantization bits globally instead of repairing critical layers — Global quantization granularity — What does it mean?

A 9-model causal study finds quantization damage is diffuse: at a small matched budget, a finer group size beats repairing a few layers.

LLM

Translate KV states across model families to cut prefill 67% — Cross-model KV translation — What does it mean?

A learned layer rewrites one model's KV cache into another model's format, so the next model resumes instead of prefilling: 899 to 138 ms.

LLM

vLLM 0.28 ships disk-backed tiered KV offload — Partial loads from a lower cache tier — What does it mean?

A tier below the GPU can hand back part of a cached prefix. vLLM 0.28 lets that partial answer be kept rather than thrown away.

LLM

Masked diffusion serving measured at 16× batch throughput — Denoising-step batching — What does it mean?

Only 24% of a masked-diffusion request's wall clock is GPU math. Line 16 requests up on the same denoising step and throughput rises 16×.

LLM

AsymSpec gives the small drafter the full context and the large verifier a compressed one — Context-asymmetric speculative decoding — What does it mean?

Speculative decoding usually shows both models the same context. AsymSpec shows the cheap model everything and the expensive model a summary.

LLM

KeysAndValues finetunes models under the KV policy they will serve with — KV policy co-adaptation — What does it mean?

Long-context models train seeing every past token, then get served seeing a fraction of them. This paper applies the eviction policy during finetuning instead.

LLM

Show chunked prefill beats elastic KV-cache reclamation — Reclaiming the prefill reserve — What does it mean?

A serving engine holds GPU memory back for the next prefill. A clever allocator can lend it out — a smaller prefill chunk frees more.

LLM

Relation reports lower loss than MHA at 10M–100M params — Self and Exchange relations — What does it mean?

Attention normalizes pairwise scores into weights immediately. Relation splits that same evidence into Self and Exchange channels first.