AI Explained

Plain explanations of trending AI concepts, with live visualizations.

LLM

AMD serves NVFP4 checkpoints on MXFP4-only GPUs — Load-time NVFP4-to-MXFP4 requantization — What does it mean?

SGLang rewrites NVFP4 weights into MXFP4 one layer at a time as the model loads, unlocking AMD's native 4-bit path.

LLM

Compress reasoning KV caches 5.8× with beacon queries — Query-cluster KV residency prediction — What does it mean?

Beacon queries predict which distant KV pairs a reasoning trace will revisit, so recency-based eviction stops guessing.

LLM

Virtualize million-token agent workspaces across GPU, RAM, and NVMe — Query-dependent execution view over paged KV memory — What does it mean?

Keep the whole agent history as paged KV, then load only the blocks this step needs into the native window.

LLM

Save 25% training FLOPs with tuned layer dropout — Progressive layer dropout for depth-elastic transformers — What does it mean?

Switching whole transformer blocks off during training saves FLOPs and leaves a model you can safely run short.

LLM

A serving study traces agent irreproducibility to prefix-cache state — Cache-state divergence — What does it mean?

Prefix caching made 36.2% of identical agent runs diverge, 75.0% at 4-bit, because the request never records the cache's state.

LLM

Tencent open-sources Hy4 Preview at 770B — Identity Hyper-Connections — What does it mean?

Hy4 Preview carries four parallel residual streams between layers instead of the transformer's usual single shared one.

LLM

Tencent open-sources Hy4 Preview at 770B — IndexCache cross-layer index reuse — What does it mean?

Hy4 Preview builds its sparse-attention index once and reuses it across layers, instead of having every attention layer rebuild it.

LLM

TensorRT-LLM makes KV cache manager V2 the default — Distributed KV pool rebalancing — What does it mean?

KV memory is split into pools by a guess made at startup. V2 moves the line between them at runtime, without breaking a captured CUDA graph.

LLM

SMELT loops the middle half of an MoE transformer twice — Looped depth reuse — What does it mean?

Depth from running the same middle layers twice, not from adding new ones — at matched FLOPs, parameters and KV cache.

LLM

Nemotron-H 8B pretrains in FP4 with no Hadamard transform — UE5M3 block scaling — What does it mean?

The FP4 payload gets the headlines; the format of the shared block scale decides whether 4-bit training works.

LLM

A 3x speedup for long reasoning, with no retraining — Prefix Sliding — What does it mean?

Keep the prefix and the last few thousand reasoning tokens, drop the middle: memory stops growing with the length of the thought.

LLM

Route local agent inference across mixed devices with NVIDIA PAIR — Request-level routing — What does it mean?

PAIR routes each whole inference request to one local device instead of splitting it, so more agent jobs run at once, not faster.