AI Explained
Plain explanations of trending AI concepts, with live visualizations.
TokenStack moves hot KV blocks into HBM-PIM — Processing-in-memory attention — What does it mean?
Most KV-cache work shrinks the bytes so fewer of them travel. TokenStack instead puts attention hardware on the memory layers the hot blocks already sit on.
RMM cuts Transformer matmuls without touching the weights — Contraction-dimension slicing — What does it mean?
Most inference savings shrink the numbers or skip whole layers. RMM keeps the weights and drops slices of the dimension a matmul sums over.
Relation reports lower loss than MHA at 10M–100M params — Self and Exchange relations — What does it mean?
Attention normalizes pairwise scores into weights immediately. Relation splits that same evidence into Self and Exchange channels first.
Heal compressed 4-bit LLMs by distilling from the original — Quantization-Aware Healing — What does it mean?
A compress-then-quantize pipeline leaves two possible teachers. QAH picks the original model, not the shrunken checkpoint in the middle.
Intel reveals a 480GB air-cooled inference GPU — Capacity-first inference memory — What does it mean?
Almost every inference accelerator buys bandwidth with HBM. Crescent Island reaches for capacity instead, with 480GB of slower LPDDR5X.
CacheRoute sends repeated prefixes to the server that still holds them — Prefix-affinity routing — What does it mean?
A prefix cache only pays off if the request lands on the machine that has it. CacheRoute makes that a routing decision, and plans it.
ReCache reuses tool-schema KV blocks across agent calls — Composition-invariant KV blocks — What does it mean?
Prefix caching only pays if the prompt matches token for token. ReCache encodes each tool schema on its own, so any order still hits.
Compress CoT reasoning with reusable memory scaffolds — Context-Generation Substitution Law — What does it mean?
Generated reasoning tokens are the expensive kind. This paper puts the reasoning in the prompt instead, and names the exchange rate.
AsmEvo tunes AMD GPU kernels at the assembly level — Correctness-gated assembly optimization — What does it mean?
AsmEvo lets an agent rewrite AMDGPU assembly, keeping only the edits whose output still matches the original on identical launches.
Ring-Zero scales zero-RL reasoning to 1T parameters — Trillion-scale zero-RL — What does it mean?
Ring-Zero runs zero-RL — reward-only RL with no SFT warm-up — at 1T parameters and finds a two-phase discovery-then-sharpening learning pattern.
MemOps benchmarks agent memory as lifecycle operations — Memory lifecycle operations — What does it mean?
MemOps grades an agent's long-term memory as four operations — remember, forget, update, reflect — not one static store of facts.
Agent optimizer study shows regression control compounds gains — Regression control in continual optimization — What does it mean?
An agent optimizer keeps compounding gains only when regression control steers it past shortcuts that boost new tasks by erasing old wins.











