AI Explained

Plain explanations of trending AI concepts, with live visualizations.

Agent

Maestro paper — RL orchestrator over frozen experts — What does it mean?

Maestro trains a 4B RL policy that picks (expert, skill) per task from a frozen model pool — 70.1% across 10 multimodal benchmarks, beating GPT-5 (69.3%), and routes to new experts without retraining.

Agent

Boiling the Frog paper — Multi-turn norm erosion vs single-prompt agent safety — What does it mean?

Boiling the Frog walks agents from benign edits to risk-bearing actions across multiple turns — averaging 44.4% attack success on 9 frontier agents that would refuse the same final message asked alone.

LLM

Gated DeltaNet-2 paper — Decoupled channel-wise erase/write gates — What does it mean?

Gated DeltaNet-2 replaces the single scalar gate of prior linear-attention models with two independent per-channel vectors — one for erasing old state, one for writing new state — so each dimension of the recurrent memory can decay at its own rate.

Agent

Camouflage Injection paper — Camouflage Detection Gap — What does it mean?

Injection payloads rewritten in a document's own domain vocabulary slip past current detectors: Llama 3.1 8B drops from 93.8% to 9.7% caught, Llama Guard 3 catches zero.

LLM

ACC paper — Tool-output unmasking — What does it mean?

ACC reformats agent trajectories as long-context QA pairs by unmasking tool outputs — Qwen3-30B-A3B gains +18.1 MRCR, matching Qwen3-235B-A22B 8× smaller.

GPU

NVIDIA Vera Rubin NVL72 — Rack-scale NVLink domain — What does it mean?

Vera Rubin NVL72 wires all 72 Rubin GPUs in a rack into one sixth-gen NVLink domain — collectives no longer cross PCIe-over-network between 8-GPU islands.

LLM

RELEX paper — Rank-1 RLVR weight-trajectory extrapolation — What does it mean?

RELEX exploits the empirical finding that RLVR fine-tuning weight trajectories are near rank-1 — fit a line through 15% of training steps and extrapolate; matches full RLVR quality from a fraction of the compute.

LLM

OScaR paper — Token Norm Imbalance — What does it mean?

Token Norm Imbalance — a few tokens carry outsized KV norms along the sequence axis. Channel rotation can't flatten them; OScaR can.

LLM

MSSP paper — Scale-stable parameterization beyond muP — What does it mean?

MSSP applies Dynamical Mean Field Theory to MoE training and derives a parameterization that — unlike muP — keeps the optimal learning rate stable as both model width and expert count scale.

LLM

Mix-Quant paper — NVFP4 prefill + BF16 decode — What does it mean?

Mix-Quant quantizes only the prefill phase to NVFP4 and keeps decode in BF16 — up to 3× prefill speedup with task performance largely preserved, because the two phases sit on opposite sides of the roofline.

LLM

PSD paper — Parallel speculative decoding for diffusion LLMs — What does it mean?

Parallel speculative decoding gets up to 5.5× tokens per forward pass on diffusion LLMs by attacking spatial and temporal axes at once.

Agent

OpenComputer paper — Verifier-grounded benchmark synthesis — What does it mean?

OpenComputer builds 1,000 computer-use tasks across 33 desktop apps by writing the executable verifier first, calibrating it, then synthesizing tasks that ground into the verifier endpoints — GPT-5.4 hits 68.3% while open-source agents collapse to as low as 5.7%.