AI Explained

Plain explanations of trending AI concepts, with live visualizations.

Agent

TokenCast — Segment-level token forecasting — What does it mean?

An agent re-reads its whole context on every call, so TokenCast forecasts spend by segment: direct cost plus growth later calls pay for.

Agent

Failure-Transparent Agents — Structured evidence contracts — What does it mean?

After a tool fails, make the agent fill STATUS, EVIDENCE, LIMITATION and NEXT ACTION; false success claims were 0.8%, not 22.8%.

LLM

CoWindow Attention splits distant history across heads — Complementary per-head windows — What does it mean?

Each KV head keeps the recent window but watches only its own slice of the far past; together the heads still see every token.

LLM

SANTA++ uses 16–22% of dense KV reads — Inverse-probability sampled attention — What does it mean?

SANTA++ samples groups of cached keys and reweights each by 1/p, estimating full attention from 16–22% of the KV reads.

LLM

KV-streams — Keeping the KV cache through context compaction — What does it mean?

Delete old KV-cache entries in place instead of starting a fresh trace, so compaction never re-prefills the kept context.

LLM

Allocate low-bit recurrent state precision by channel decay — Decay-aware bit allocation for recurrent states — What does it mean?

A rounding error in a recurrent state lasts as long as its row remembers, so slowly decaying rows get more bits at the same average.

Agent

EffectMatch checks what an agent's tool call actually changed — Effect-based authorization — What does it mean?

An approved tool call can succeed and still leave a change nobody approved. EffectMatch checks the real changes before commit.

Agent

Where Does Exactly-Once Live? — Idempotency keys vs read-back verification for exactly-once tool effects — What does it mean?

A timed-out agent write may still land. Read-back checks miss it; idempotency keys cut duplicates from 28% to 4% in the LIMBO benchmark.

LLM

CacheReforge — Stale KV cache repair after LoRA adapter updates — What does it mean?

A LoRA update makes cached K/V stale. CacheReforge recomputes about 5% of layers instead of half and removes 92% of the output drift.

LLM

Reparameterize output heads before W4 quantization — Softmax-preserving head shift — What does it mean?

Softmax ignores shared logit shifts, so choosing an equivalent output head before rounding can make a 4-bit head far more faithful.

LLM

Crossflow — Revocable decode-node leases — What does it mean?

Let decode GPUs lend short-lived, revocable capacity to prefill when prompt demand spikes, instead of re-splitting a fixed P/D cluster.

LLM

From Memory Budgets to Risk Targets — Risk-controlled KV cache eviction — What does it mean?

Pick how much KV cache to evict by certifying how often requests break, not by average score, and keep the full cache if nothing passes.