AI Explained

Plain explanations of trending AI concepts, with live visualizations.

Agent

RTK reported 89% token savings and DeepSeek's cost rose 17% — Turn amplification — What does it mean?

RTK reported 89% token savings. DeepSeek's per-task cost still rose 17%, because thinner turns bought more turns.

Agent

RAG-Safety-Bench — Benign-context safety degradation — What does it mean?

For some models, harmless retrieved documents lead to unsafe generation — RAG-Safety-Bench isolates that from retriever quality.

Agent

AgentZip compresses agent sandboxes 8.7× — Cross-sandbox memory redundancy — What does it mean?

A compression ratio is decided by what you compare a page against — and clones of one template are the best comparison you have.

Agent

AgentZip compresses agent sandboxes 8.7× — LLM-wait latency hiding — What does it mean?

Do the expensive sandbox work in the window where the agent is already waiting on the model, and most of its cost stops being visible.

LLM

Compress context up to 266× with one ratio-adaptive model — Matryoshka memory budgets — What does it mean?

One soft-context compressor whose every prefix is a valid coarser summary, so each input picks its own compression ratio.

LLM

Make NVMe KV-cache loads 2x faster with scheduler-aware preloading — External KV-cache break-even — What does it mean?

An external KV-cache hit is an admission decision with a break-even, not a free lookup.

Agent

The model with the higher token price cost about 23% less per passing answer — Cost per successful outcome — What does it mean?

Divide all spend, failures included, by the outcomes you keep — and the per-token ranking can invert.

GPU

Decoupling prefill and decode power saves 32.3% of a lane pair's electricity — Phase-decoupled power control — What does it mean?

Prefill and decode sit on opposite sides of the roofline, so one power setting for both is wrong for one. Give each lane its own rule.

LLM

Audit finds a typical DeepSeek-V4-Flash site uses about two of its four residual streams — Near-identity late residual mixing — What does it mean?

DeepSeek-V4-Flash carries four residual streams, but across layers 22–42 it barely mixes them.

LLM

Pretrained LLMs resist 4-bit quantization partly because layer errors cancel — Counteracting quantization error — What does it mean?

Pretrained LLMs survive 4-bit rounding because each layer's new error opposes the one it inherited — a cancellation learned in pretraining.

LLM

Cut RAG TTFT 80% with selectively recomputed KV-cache chunks — Selective KV recomputation — What does it mean?

Reusing precomputed KV caches per RAG chunk is fast but drops cross-chunk attention; re-running a chosen subset buys the quality back.

GPU

Replace exp-then-quantize softmax to cut vector latency 40.33% — Exp-free softmax quantization to E2M1 probability codes — What does it mean?

EFQ-Softmax emits 4-bit E2M1 probabilities directly from attention scores, deleting the exp() stage inside FlashAttention.