AI Explained

Plain explanations of trending AI concepts, with live visualizations.

LLM

ARMT extends LLM context with recurrent associative memory — Associative recurrent memory — What does it mean?

ARMT extends an LLM's context with a fixed-size recurrent associative memory instead of a growing KV cache — flat memory as text grows, ~30% fewer FLOPs.

LLM

Self-guided TTT adapts long-context LLMs on selected spans — Evidence-span test-time training — What does it mean?

Before answering, the model highlights the evidence spans that matter and briefly trains on only those — a per-question tune-up that recovers long-context accuracy.

LLM

KV-PRM scores agent rollouts from the KV cache — Verify-token scoring against the KV cache — What does it mean?

KV-PRM grades a reasoning trajectory from the KV cache the model already built — one verify token, not a full re-encode of the whole trace.

LLM

vLLM 0.25 makes Model Runner V2 the dense default — Retiring PagedAttention — What does it mean?

vLLM 0.25 removes PagedAttention and makes Model Runner V2 the dense default, and reports full CUDA graphs for decode.

LLM

Linearization study isolates cache routing for long-context attention — Training-free linear attention — What does it mean?

A frozen model is converted to linear attention without retraining — sink tokens, a short convolution, and fixed-budget cache routing add back the long-context recall a plain swap loses.

LLM

Quantization study finds accuracy can hide behavior drift — Correctness agreement — What does it mean?

A quantized LLM can keep the same accuracy yet change which answers it gets right — correctness agreement measures that hidden drift.

LLM

NVIDIA Audex — Unified audio-text token space — What does it mean?

Audex is one text decoder that also hears and speaks: audio shares the token space of text, with little to no drop in text skill.

LLM

MAESTRO prunes MoE experts with routing-aware Markov chains — Markov-chain expert pruning — What does it mean?

MAESTRO models MoE routing as a Markov chain, ranks experts by their long-run share of traffic, and prunes the rarely-visited ones — keeping up to 10.61% more performance at 50% compression.

LLM

DominoTree speeds speculative decoding with conditional draft trees — Conditional draft-tree scoring — What does it mean?

DominoTree scores speculative-decoding draft trees by how well each token follows the last, not one token at a time — up to 6.6× on Qwen3-4B.

LLM

TF-Engram uses SSD-backed phrase memory for train-free LLM recall — Predictive prefetching that hides SSD latency — What does it mean?

TF-Engram stores phrase memories across GPU-DRAM-SSD and prefetches them so a slow SSD read overlaps decode instead of stalling it.

LLM

DeLS-Spec adds short-context heads to block drafting — Long-short logit fusion — What does it mean?

DeLS-Spec fuses a short-context head into a block drafter, discounting common words, so a fast block draft finally picks up the local context its parallel guesses skipped.

LLM

SIS turns off-policy RL tokens into on-policy updates — Selective Importance Sampling — What does it mean?

Selective Importance Sampling accepts the agreeing tokens as on-policy (ratio 1), taming the variance when RL reuses off-policy rollouts.