AI Explained
Plain explanations of trending AI concepts, with live visualizations.
GATS plans agent tasks with zero LLM calls during the search — World-model tree search — What does it mean?
GATS plans an agent's actions by searching a learned world model — the search makes zero LLM calls, and it solved every task in the paper's test.
FastTPS fuses token-phase LLM inference on AI accelerators — Reloading-free KV-cache concatenation — What does it mean?
FastTPS runs LLM decode without recopying the KV cache to append: tiled attention then hits 93% peak memory bandwidth, 6× faster on an NPU.
ARMT extends LLM context with recurrent associative memory — Associative recurrent memory — What does it mean?
ARMT extends an LLM's context with a fixed-size recurrent associative memory instead of a growing KV cache — flat memory as text grows, ~30% fewer FLOPs.
Self-guided TTT adapts long-context LLMs on selected spans — Evidence-span test-time training — What does it mean?
Before answering, the model highlights the evidence spans that matter and briefly trains on only those — a per-question tune-up that recovers long-context accuracy.
NVIDIA frames Vera CPU around fast agent steps — The agent-step CPU bottleneck — What does it mean?
NVIDIA argues an agent step spends real time on serial CPU work between GPU calls, so Vera targets max single-thread speed to keep GPUs from idling.
KV-PRM scores agent rollouts from the KV cache — Verify-token scoring against the KV cache — What does it mean?
KV-PRM grades a reasoning trajectory from the KV cache the model already built — one verify token, not a full re-encode of the whole trace.
Danus coordinates math agents with fact-graph memory — Verifier-gated fact graph — What does it mean?
Danus lets parallel math agents build long proofs by gating every claim through a verifier into one shared fact graph.
Agora auctions each agent reasoning step to expert models — Auction-based task allocation — What does it mean?
Agora puts each reasoning step up for auction so expert models bid by rectified competence and cost — one dial slides the whole system along the cost-quality frontier.
vLLM 0.25 makes Model Runner V2 the dense default — Retiring PagedAttention — What does it mean?
vLLM 0.25 removes PagedAttention and makes Model Runner V2 the dense default, and reports full CUDA graphs for decode.
Linearization study isolates cache routing for long-context attention — Training-free linear attention — What does it mean?
A frozen model is converted to linear attention without retraining — sink tokens, a short convolution, and fixed-budget cache routing add back the long-context recall a plain swap loses.
STRACE extracts causal root causes from noisy agent traces — Causal trace localization — What does it mean?
STRACE shrinks a failed agent's huge run log down to the few steps that actually caused the failure — causal localization, not length-based trimming — lifting a verifier agent 42.5% → 58.5%.
Microsoft open-sources Flint for agent-generated visualizations — Semantic intermediate representation — What does it mean?
Microsoft's Flint has the agent write a short chart spec; a compiler fills in the brittle low-level details — for more reliable agent-drawn charts.











