AI Explained
Plain explanations of trending AI concepts, with live visualizations.
ARMT extends LLM context with recurrent associative memory — Associative recurrent memory — What does it mean?
ARMT extends an LLM's context with a fixed-size recurrent associative memory instead of a growing KV cache — flat memory as text grows, ~30% fewer FLOPs.
Self-guided TTT adapts long-context LLMs on selected spans — Evidence-span test-time training — What does it mean?
Before answering, the model highlights the evidence spans that matter and briefly trains on only those — a per-question tune-up that recovers long-context accuracy.
KV-PRM scores agent rollouts from the KV cache — Verify-token scoring against the KV cache — What does it mean?
KV-PRM grades a reasoning trajectory from the KV cache the model already built — one verify token, not a full re-encode of the whole trace.
vLLM 0.25 makes Model Runner V2 the dense default — Retiring PagedAttention — What does it mean?
vLLM 0.25 removes PagedAttention and makes Model Runner V2 the dense default, and reports full CUDA graphs for decode.
Linearization study isolates cache routing for long-context attention — Training-free linear attention — What does it mean?
A frozen model is converted to linear attention without retraining — sink tokens, a short convolution, and fixed-budget cache routing add back the long-context recall a plain swap loses.
Quantization study finds accuracy can hide behavior drift — Correctness agreement — What does it mean?
A quantized LLM can keep the same accuracy yet change which answers it gets right — correctness agreement measures that hidden drift.
NVIDIA Audex — Unified audio-text token space — What does it mean?
Audex is one text decoder that also hears and speaks: audio shares the token space of text, with little to no drop in text skill.
MAESTRO prunes MoE experts with routing-aware Markov chains — Markov-chain expert pruning — What does it mean?
MAESTRO models MoE routing as a Markov chain, ranks experts by their long-run share of traffic, and prunes the rarely-visited ones — keeping up to 10.61% more performance at 50% compression.
DominoTree speeds speculative decoding with conditional draft trees — Conditional draft-tree scoring — What does it mean?
DominoTree scores speculative-decoding draft trees by how well each token follows the last, not one token at a time — up to 6.6× on Qwen3-4B.
TF-Engram uses SSD-backed phrase memory for train-free LLM recall — Predictive prefetching that hides SSD latency — What does it mean?
TF-Engram stores phrase memories across GPU-DRAM-SSD and prefetches them so a slow SSD read overlaps decode instead of stalling it.
DeLS-Spec adds short-context heads to block drafting — Long-short logit fusion — What does it mean?
DeLS-Spec fuses a short-context head into a block drafter, discounting common words, so a fast block draft finally picks up the local context its parallel guesses skipped.
SIS turns off-policy RL tokens into on-policy updates — Selective Importance Sampling — What does it mean?
Selective Importance Sampling accepts the agreeing tokens as on-policy (ratio 1), taming the variance when RL reuses off-policy rollouts.











