AI Explained

Plain explanations of trending AI concepts, with live visualizations.

LLM

SwitchSD reads copy intent from the model's own states — Copy-intent probe drafting-mode switch — What does it mean?

SwitchSD reads copy intent from the target model's hidden states, switching to cheap copying only when the probe predicts real repetition.

LLM

NVIDIA replaces GenAI-Perf with multiprocess AIPerf — Multiprocess load generation — What does it mean?

A single-process benchmark client is capped by Python's GIL, so the throughput you measure can be the client's ceiling, not the server's.

LLM

Let models trigger full attention only when recall helps — On-Demand Attention recall-head gating — What does it mean?

A small recall head decides, per decode step, whether reading the full KV cache is worth it — at long context, skipping it cuts decode cost.

LLM

DeepSeek-V4.1-Flash runs prefill on 8B parameters and decode on 16B — Causal Encoder-Decoder — What does it mean?

DeepSeek-V4.1-Flash runs prefill on an 8B parameter path and decode on a 16B one — a split in the parameters, not in the machines.

Agent

Chronicle replays agent incidents as CI tests — Cut-point replay vs stubbing every boundary — What does it mean?

Freeze the boundaries you are not testing, run the one you are live, and an incident becomes a CI test.

LLM

D-Quant drifts variable-length KV codes into fixed-size token streams — Entropy-coded KV cache — What does it mean?

Short codewords for common KV values, long for rare ones, then drifted back into fixed-size per-token streams.

Agent

Magenta closes informal math reasoning with Lean-guided repair — Error-attribution routing — What does it mean?

Magenta names why a Lean proof failed, then routes the retry to a re-derivation or a one-step patch instead of resampling.

Agent

TypeSafe releases Jev — Typed parallel decisions — What does it mean?

Jev fixes the output type before it runs, samples every field in parallel, and attaches a calibrated probability to each one.

LLM

JustFit serves a 212,992-token context on a 24 GiB laptop — Just-in-time state residency — What does it mean?

JustFit holds a 212,992-token context on a 24 GiB laptop by scheduling when each piece of state is resident, not by shrinking the weights.

GPU

Devin-built GPU sieving factored RSA-260 — Preemptible sieving vs the interconnect-bound solve — What does it mean?

Sieving survives on preemptible GPUs because losing a unit costs milliseconds. The block Wiedemann solve cannot, so only it pays for NVLink.

Agent

Emergence World runs 16-day agent worlds — Non-compositional alignment — What does it mean?

Eight 10-agent worlds, 16 days, no survivors: agents detected the injection and still acted on it up to 46 hours later.

Agent

GAUGE finds 57.5% of agent chats rated satisfying still failed the user's task — Ranking validity vs construct validity — What does it mean?

An offline judge can rank agents right and still measure the wrong thing: 57.5% of chats rated satisfying failed the user's task.