AI Explained

Plain explanations of trending AI concepts, with live visualizations.

LLM

SIS turns off-policy RL tokens into on-policy updates — Selective Importance Sampling — What does it mean?

Selective Importance Sampling accepts the agreeing tokens as on-policy (ratio 1), taming the variance when RL reuses off-policy rollouts.

LLM

Discrete diffusion theory unifies denoisers, scores & bridge predictors — One object, three coordinates — What does it mean?

The three ways to train a diffusion language model — denoiser, score, and bridge — turn out to be one object written in different coordinates.

LLM

Direct-OPD transfers weak-model RL gains as log-ratio rewards — Weak-to-strong reward transfer — What does it mean?

Direct-OPD reuses a small model's before/after RL log-ratio as a dense reward to lift a stronger model — no RL on the big model.

LLM

LOCOS finds non-literal retrieval heads by scoring logit contribution — Logit-Contribution Scoring — What does it mean?

LOCOS finds the attention heads that answer from meaning, not copied words, by scoring how much each head writes toward the answer token.

LLM

LACUNA tests whether LLM unlearning hits the right weights — Output-level vs weight-level unlearning evaluation — What does it mean?

LACUNA plants a fact in known weights to test whether unlearning erases it, or just hides the output while a resurfacing attack revives it.

LLM

Google releases TabFM for zero-shot tabular prediction — Tabular in-context learning — What does it mean?

TabFM predicts on a spreadsheet it has never seen by reading the whole table as one prompt — in-context learning, no per-dataset training.

LLM

BlockSearch traces million-token retrieval collapse — Attention dilution — What does it mean?

BlockSearch traces long-context retrieval collapse to attention dilution — the softmax denominator drowning the gold document — and fixes it with a length-aware softmax.

LLM

Orca learns a unified world model over video and events — Next-State-Prediction — What does it mean?

Orca trains one shared world model with a single objective — predict the next state — then reads it out as text, image, or action.

LLM

CausalMix picks LLM pretraining data mixtures via causal inference — CATE-based data mixture selection — What does it mean?

CausalMix picks an LLM's pretraining data mix by causal inference: estimate each mix's causal effect from 512 tiny runs, then scale to 7B.

LLM

ELDR routes MoE decode by expert locality, cutting TPOT up to 13.9% — Expert-locality-aware decode routing — What does it mean?

ELDR predicts a request's experts and routes decode to the worker whose zone matches, so fewer experts reload — cutting TPOT 5.9–13.9%.

LLM

DOPD dodges the 'privilege illusion' — Dual on-policy distillation — What does it mean?

DOPD routes each token's distillation signal by advantage, so a small model learns real skill instead of an answer key it can never hold.

LLM

BlockPilot gives diffusion speculative decoding 4.2× — Instance-adaptive draft block sizing — What does it mean?

BlockPilot learns a per-input policy that predicts the best draft block size for diffusion speculative decoding, hitting 4.20× on a 4B model.