AI Explained
Plain explanations of trending AI concepts, with live visualizations.
Maestro paper — RL orchestrator over frozen experts — What does it mean?
Maestro trains a 4B RL policy that picks (expert, skill) per task from a frozen model pool — 70.1% across 10 multimodal benchmarks, beating GPT-5 (69.3%), and routes to new experts without retraining.
Boiling the Frog paper — Multi-turn norm erosion vs single-prompt agent safety — What does it mean?
Boiling the Frog walks agents from benign edits to risk-bearing actions across multiple turns — averaging 44.4% attack success on 9 frontier agents that would refuse the same final message asked alone.
Gated DeltaNet-2 paper — Decoupled channel-wise erase/write gates — What does it mean?
Gated DeltaNet-2 replaces the single scalar gate of prior linear-attention models with two independent per-channel vectors — one for erasing old state, one for writing new state — so each dimension of the recurrent memory can decay at its own rate.
Camouflage Injection paper — Camouflage Detection Gap — What does it mean?
Injection payloads rewritten in a document's own domain vocabulary slip past current detectors: Llama 3.1 8B drops from 93.8% to 9.7% caught, Llama Guard 3 catches zero.
ACC paper — Tool-output unmasking — What does it mean?
ACC reformats agent trajectories as long-context QA pairs by unmasking tool outputs — Qwen3-30B-A3B gains +18.1 MRCR, matching Qwen3-235B-A22B 8× smaller.
NVIDIA Vera Rubin NVL72 — Rack-scale NVLink domain — What does it mean?
Vera Rubin NVL72 wires all 72 Rubin GPUs in a rack into one sixth-gen NVLink domain — collectives no longer cross PCIe-over-network between 8-GPU islands.
RELEX paper — Rank-1 RLVR weight-trajectory extrapolation — What does it mean?
RELEX exploits the empirical finding that RLVR fine-tuning weight trajectories are near rank-1 — fit a line through 15% of training steps and extrapolate; matches full RLVR quality from a fraction of the compute.
OScaR paper — Token Norm Imbalance — What does it mean?
Token Norm Imbalance — a few tokens carry outsized KV norms along the sequence axis. Channel rotation can't flatten them; OScaR can.
MSSP paper — Scale-stable parameterization beyond muP — What does it mean?
MSSP applies Dynamical Mean Field Theory to MoE training and derives a parameterization that — unlike muP — keeps the optimal learning rate stable as both model width and expert count scale.
Mix-Quant paper — NVFP4 prefill + BF16 decode — What does it mean?
Mix-Quant quantizes only the prefill phase to NVFP4 and keeps decode in BF16 — up to 3× prefill speedup with task performance largely preserved, because the two phases sit on opposite sides of the roofline.
PSD paper — Parallel speculative decoding for diffusion LLMs — What does it mean?
Parallel speculative decoding gets up to 5.5× tokens per forward pass on diffusion LLMs by attacking spatial and temporal axes at once.
OpenComputer paper — Verifier-grounded benchmark synthesis — What does it mean?
OpenComputer builds 1,000 computer-use tasks across 33 desktop apps by writing the executable verifier first, calibrating it, then synthesizing tasks that ground into the verifier endpoints — GPT-5.4 hits 68.3% while open-source agents collapse to as low as 5.7%.