AI Explained
Plain explanations of trending AI concepts, with live visualizations.
TokenCast — Segment-level token forecasting — What does it mean?
An agent re-reads its whole context on every call, so TokenCast forecasts spend by segment: direct cost plus growth later calls pay for.
Failure-Transparent Agents — Structured evidence contracts — What does it mean?
After a tool fails, make the agent fill STATUS, EVIDENCE, LIMITATION and NEXT ACTION; false success claims were 0.8%, not 22.8%.
CoWindow Attention splits distant history across heads — Complementary per-head windows — What does it mean?
Each KV head keeps the recent window but watches only its own slice of the far past; together the heads still see every token.
SANTA++ uses 16–22% of dense KV reads — Inverse-probability sampled attention — What does it mean?
SANTA++ samples groups of cached keys and reweights each by 1/p, estimating full attention from 16–22% of the KV reads.
KV-streams — Keeping the KV cache through context compaction — What does it mean?
Delete old KV-cache entries in place instead of starting a fresh trace, so compaction never re-prefills the kept context.
Allocate low-bit recurrent state precision by channel decay — Decay-aware bit allocation for recurrent states — What does it mean?
A rounding error in a recurrent state lasts as long as its row remembers, so slowly decaying rows get more bits at the same average.
EffectMatch checks what an agent's tool call actually changed — Effect-based authorization — What does it mean?
An approved tool call can succeed and still leave a change nobody approved. EffectMatch checks the real changes before commit.
Where Does Exactly-Once Live? — Idempotency keys vs read-back verification for exactly-once tool effects — What does it mean?
A timed-out agent write may still land. Read-back checks miss it; idempotency keys cut duplicates from 28% to 4% in the LIMBO benchmark.
CacheReforge — Stale KV cache repair after LoRA adapter updates — What does it mean?
A LoRA update makes cached K/V stale. CacheReforge recomputes about 5% of layers instead of half and removes 92% of the output drift.
Reparameterize output heads before W4 quantization — Softmax-preserving head shift — What does it mean?
Softmax ignores shared logit shifts, so choosing an equivalent output head before rounding can make a 4-bit head far more faithful.
Crossflow — Revocable decode-node leases — What does it mean?
Let decode GPUs lend short-lived, revocable capacity to prefill when prompt demand spikes, instead of re-splitting a fixed P/D cluster.
From Memory Budgets to Risk Targets — Risk-controlled KV cache eviction — What does it mean?
Pick how much KV cache to evict by certifying how often requests break, not by average score, and keep the full cache if nothing passes.