AI Explained

Plain explanations of trending AI concepts, with live visualizations.

LLM

D-Cut prunes speculative decoding verification across batches — Cross-request draft pruning — What does it mean?

Speculative decoding gives every request the same draft depth. D-Cut pools the drafts from all of them and verifies only the ones likely to be accepted.

GPU

Atrex-Bench tests LLM-written kernels on production traces — Trace-weighted kernel benchmarking — What does it mean?

A kernel benchmark that samples operators and shapes from real serving traces and weights each one by the GPU time it actually burns.

Agent

Plan evaluator study exposes omission incentives — Deletion non-monotonicity — What does it mean?

Deleting a step a plan needs can raise its score — the scorer drops both the step's cost and its chance of failing.

Agent

Long-Horizon-Terminal-Bench grades agent progress densely — Partial-reward threshold — What does it mean?

Pass/fail turns a 231-episode run into one bit. A partial-reward threshold grades how far the agent got, and moving the line moves the score.

LLM

LLM-as-judge bias appears as activation geometry — Steerable bias directions — What does it mean?

LLM-as-judge bias is a steerable direction in the judge's hidden state — a study that reads it, predicts it, and steers it back to baseline.

Agent

OpenAI trains GPT-Red to harden agents against prompt injection — Self-play red-teaming — What does it mean?

OpenAI trained an attacker model to invent prompt injections, then fed everything it found back into the defender's training — a red-teamer that never ships.

GPU

vLLM 0.25.1 stops a fused kernel from corrupting NVFP4 models — Mixed-dtype quant-fusion guard — What does it mean?

vLLM 0.25.1 adds a dtype check so a fused allreduce+RMSNorm+quant kernel stops silently corrupting mixed-precision NVFP4 models.

Agent

E3 cuts coding-agent scope before expanding context — Estimate-Execute-Expand — What does it mean?

E3 is a coding-agent strategy: start from the smallest scope, expand only when a check fails, cutting cost ~85% at the same success.

LLM

AVQ-Attention refines codewords where attention mass concentrates — Adaptive vector-quantized attention — What does it mean?

AVQ-Attention represents keys as a few codewords, then adds detail only where attention concentrates — turning O(N²) attention into O(MN).

LLM

KronQ adds gradient covariance to LLM quantization — Kronecker-factored Hessian — What does it mean?

KronQ scores the rounding error by its output impact, not just its input size — so 2-bit LLaMA-3-70B holds instead of collapsing.

Agent

Interaction scaling grounds agent feedback loops — Instrument-grounded feedback loops — What does it mean?

An agent keeps improving when it acts, measures the result with a tool, and revises — not by thinking longer or trusting a review by eye.

LLM

HCRMap places hot MoE experts across chiplet memory tiers — Hotness-aware MoE expert replica placement — What does it mean?

HCRMap ranks each MoE expert by how hot it runs, then promotes or evicts its copies across memory tiers, cutting latency up to 46.7%.