AI Explained

Plain explanations of trending AI concepts, with live visualizations.

Agent

GATS plans agent tasks with zero LLM calls during the search — World-model tree search — What does it mean?

GATS plans an agent's actions by searching a learned world model — the search makes zero LLM calls, and it solved every task in the paper's test.

GPU

FastTPS fuses token-phase LLM inference on AI accelerators — Reloading-free KV-cache concatenation — What does it mean?

FastTPS runs LLM decode without recopying the KV cache to append: tiled attention then hits 93% peak memory bandwidth, 6× faster on an NPU.

LLM

ARMT extends LLM context with recurrent associative memory — Associative recurrent memory — What does it mean?

ARMT extends an LLM's context with a fixed-size recurrent associative memory instead of a growing KV cache — flat memory as text grows, ~30% fewer FLOPs.

LLM

Self-guided TTT adapts long-context LLMs on selected spans — Evidence-span test-time training — What does it mean?

Before answering, the model highlights the evidence spans that matter and briefly trains on only those — a per-question tune-up that recovers long-context accuracy.

GPU

NVIDIA frames Vera CPU around fast agent steps — The agent-step CPU bottleneck — What does it mean?

NVIDIA argues an agent step spends real time on serial CPU work between GPU calls, so Vera targets max single-thread speed to keep GPUs from idling.

LLM

KV-PRM scores agent rollouts from the KV cache — Verify-token scoring against the KV cache — What does it mean?

KV-PRM grades a reasoning trajectory from the KV cache the model already built — one verify token, not a full re-encode of the whole trace.

Agent

Danus coordinates math agents with fact-graph memory — Verifier-gated fact graph — What does it mean?

Danus lets parallel math agents build long proofs by gating every claim through a verifier into one shared fact graph.

Agent

Agora auctions each agent reasoning step to expert models — Auction-based task allocation — What does it mean?

Agora puts each reasoning step up for auction so expert models bid by rectified competence and cost — one dial slides the whole system along the cost-quality frontier.

LLM

vLLM 0.25 makes Model Runner V2 the dense default — Retiring PagedAttention — What does it mean?

vLLM 0.25 removes PagedAttention and makes Model Runner V2 the dense default, and reports full CUDA graphs for decode.

LLM

Linearization study isolates cache routing for long-context attention — Training-free linear attention — What does it mean?

A frozen model is converted to linear attention without retraining — sink tokens, a short convolution, and fixed-budget cache routing add back the long-context recall a plain swap loses.

Agent

STRACE extracts causal root causes from noisy agent traces — Causal trace localization — What does it mean?

STRACE shrinks a failed agent's huge run log down to the few steps that actually caused the failure — causal localization, not length-based trimming — lifting a verifier agent 42.5% → 58.5%.

Agent

Microsoft open-sources Flint for agent-generated visualizations — Semantic intermediate representation — What does it mean?

Microsoft's Flint has the agent write a short chart spec; a compiler fills in the brittle low-level details — for more reliable agent-drawn charts.