AI Explained
Plain explanations of trending AI concepts, with live visualizations.
NVIDIA opens two Rust paths for CUDA kernels — SIMT thread model vs tile-level programming — What does it mean?
NVIDIA opened two Rust paths to CUDA kernels, one per-thread and one per-tile. The split contrasts two established GPU programming models.
AMD serves NVFP4 checkpoints on MXFP4-only GPUs — Load-time NVFP4-to-MXFP4 requantization — What does it mean?
SGLang rewrites NVFP4 weights into MXFP4 one layer at a time as the model loads, unlocking AMD's native 4-bit path.
Compress reasoning KV caches 5.8× with beacon queries — Query-cluster KV residency prediction — What does it mean?
Beacon queries predict which distant KV pairs a reasoning trace will revisit, so recency-based eviction stops guessing.
Virtualize million-token agent workspaces across GPU, RAM, and NVMe — Query-dependent execution view over paged KV memory — What does it mean?
Keep the whole agent history as paged KV, then load only the blocks this step needs into the native window.
Save 25% training FLOPs with tuned layer dropout — Progressive layer dropout for depth-elastic transformers — What does it mean?
Switching whole transformer blocks off during training saves FLOPs and leaves a model you can safely run short.
Gate risky coding-agent actions with draft-model uncertainty — Speculative Uncertainty — What does it mean?
A small draft model reads the agent's finished trajectory once and scores how risky it looks, before anything runs.
A serving study traces agent irreproducibility to prefix-cache state — Cache-state divergence — What does it mean?
Prefix caching made 36.2% of identical agent runs diverge, 75.0% at 4-bit, because the request never records the cache's state.
LoopArena benchmarks the model that supervises a coding agent — Slice evaluation as a rank-preserving proxy — What does it mean?
Rank models on a cheap slice of the task, then check the slice ordering against the full run.
LoopArena benchmarks the model that supervises a coding agent — Controller-worker separation — What does it mean?
Hold the coding agent fixed and score only the model that decides what it does next.
Tencent open-sources Hy4 Preview at 770B — Identity Hyper-Connections — What does it mean?
Hy4 Preview carries four parallel residual streams between layers instead of the transformer's usual single shared one.
Tencent open-sources Hy4 Preview at 770B — IndexCache cross-layer index reuse — What does it mean?
Hy4 Preview builds its sparse-attention index once and reuses it across layers, instead of having every attention layer rebuild it.
TensorRT-LLM makes KV cache manager V2 the default — Distributed KV pool rebalancing — What does it mean?
KV memory is split into pools by a guess made at startup. V2 moves the line between them at runtime, without breaking a captured CUDA graph.