AI Explained

Plain explanations of trending AI concepts, with live visualizations.

GPU

NVIDIA opens two Rust paths for CUDA kernels — SIMT thread model vs tile-level programming — What does it mean?

NVIDIA opened two Rust paths to CUDA kernels, one per-thread and one per-tile. The split contrasts two established GPU programming models.

LLM

AMD serves NVFP4 checkpoints on MXFP4-only GPUs — Load-time NVFP4-to-MXFP4 requantization — What does it mean?

SGLang rewrites NVFP4 weights into MXFP4 one layer at a time as the model loads, unlocking AMD's native 4-bit path.

LLM

Compress reasoning KV caches 5.8× with beacon queries — Query-cluster KV residency prediction — What does it mean?

Beacon queries predict which distant KV pairs a reasoning trace will revisit, so recency-based eviction stops guessing.

LLM

Virtualize million-token agent workspaces across GPU, RAM, and NVMe — Query-dependent execution view over paged KV memory — What does it mean?

Keep the whole agent history as paged KV, then load only the blocks this step needs into the native window.

LLM

Save 25% training FLOPs with tuned layer dropout — Progressive layer dropout for depth-elastic transformers — What does it mean?

Switching whole transformer blocks off during training saves FLOPs and leaves a model you can safely run short.

Agent

Gate risky coding-agent actions with draft-model uncertainty — Speculative Uncertainty — What does it mean?

A small draft model reads the agent's finished trajectory once and scores how risky it looks, before anything runs.

LLM

A serving study traces agent irreproducibility to prefix-cache state — Cache-state divergence — What does it mean?

Prefix caching made 36.2% of identical agent runs diverge, 75.0% at 4-bit, because the request never records the cache's state.

Agent

LoopArena benchmarks the model that supervises a coding agent — Slice evaluation as a rank-preserving proxy — What does it mean?

Rank models on a cheap slice of the task, then check the slice ordering against the full run.

Agent

LoopArena benchmarks the model that supervises a coding agent — Controller-worker separation — What does it mean?

Hold the coding agent fixed and score only the model that decides what it does next.

LLM

Tencent open-sources Hy4 Preview at 770B — Identity Hyper-Connections — What does it mean?

Hy4 Preview carries four parallel residual streams between layers instead of the transformer's usual single shared one.

LLM

Tencent open-sources Hy4 Preview at 770B — IndexCache cross-layer index reuse — What does it mean?

Hy4 Preview builds its sparse-attention index once and reuses it across layers, instead of having every attention layer rebuild it.

LLM

TensorRT-LLM makes KV cache manager V2 the default — Distributed KV pool rebalancing — What does it mean?

KV memory is split into pools by a guess made at startup. V2 moves the line between them at runtime, without breaking a captured CUDA graph.