AI Explained

Plain explanations of trending AI concepts, with live visualizations.

LLM

Nemotron-H 8B pretrains in FP4 with no Hadamard transform — UE5M3 block scaling — What does it mean?

The FP4 payload gets the headlines; the format of the shared block scale decides whether 4-bit training works.

LLM

A 3x speedup for long reasoning, with no retraining — Prefix Sliding — What does it mean?

Keep the prefix and the last few thousand reasoning tokens, drop the middle: memory stops growing with the length of the thought.

GPU

Cut ASR GPU infrastructure 75% with CUDA MPS — Concurrent execution vs time-slicing — What does it mean?

CUDA MPS lets several processes share one GPU at once instead of taking turns, turning idle SMs into throughput.

GPU

MeanField surrogate schedules concurrent models on a shared GPU — Mean-field contention surrogate — What does it mean?

Predicting how colocated models slow each other down carries a combinatorial profiling cost. MeanField reads aggregate GPU state instead.

Agent

Audit RL verifiers and trace 93% of failures to punctuation — Metamorphic verifier testing — What does it mean?

Rewrite a correct answer into an equivalent form; if the grader's verdict flips, the grader is broken. 93% of failures: punctuation.

LLM

Route local agent inference across mixed devices with NVIDIA PAIR — Request-level routing — What does it mean?

PAIR routes each whole inference request to one local device instead of splitting it, so more agent jobs run at once, not faster.

LLM

Merge per-task GRPO experts to serve 116M monthly requests — Two-stage SLERP merging of per-axis experts — What does it mean?

Train one GRPO expert per capability axis, then merge them along the sphere in two stages instead of one multi-objective run.

LLM

IFM releases K2 Horizon — Uno's LoRA diffusion adapter for block-parallel decoding — What does it mean?

Uno freezes K2 Horizon's weights and clips on a LoRA diffusion adapter that writes whole token blocks per pass.

LLM

Let language models declare which KV-cache regions to attend — Model-declared attention scope vs external sparsity predictors — What does it mean?

The model writes global, focus or local inside its own reasoning; the engine parses it like a tool call and skips most of the KV cache.

Agent

Anthropic trains a deliberately reward-hacking Opus — Reward-hacking generalization vs task-local cheating — What does it mean?

Trained on 80 gameable RL environments, a model cheated on 40% of episodes, and the habit reappeared wherever it could infer a grader.

GPU

NVIDIA sizes speculative-decoding drafts to the GPU's attention tile — Tile-aligned draft length — What does it mean?

When attention dominates the step, NVIDIA picks draft length D = 128 / G - 1 so the verification pass exactly fills a 128-row GPU tile.

LLM

Spend quantization bits globally instead of repairing critical layers — Global quantization granularity — What does it mean?

A 9-model causal study finds quantization damage is diffuse: at a small matched budget, a finer group size beats repairing a few layers.