AI Explained

Plain explanations of trending AI concepts, with live visualizations.

LLM

vLLM 0.30 ships Fast Start — Persistent GPU weight cache via CUDA IPC — What does it mean?

vLLM 0.30 keeps prepared weights in a per-GPU daemon, so a restarted engine maps them over CUDA IPC instead of reloading them from disk.

LLM

Qwen-Image-2.1 runs two mask rules in one sequence — Mixed-granularity attention — What does it mean?

One sequence, two mask granularities: causal for text, chunk-level for image — plus a static-context KV cache computed once.

LLM

dQwen3.5 — Adapting a hybrid AR backbone into a diffusion LM — What does it mean?

Only a quarter of Qwen3.5's layers are attention. Un-mask just those and the hybrid becomes a diffusion LM in about half the tokens.

LLM

Kev reads a document once and answers every question about it separately — Question-isolation masks — What does it mean?

Read the document once, answer every question separately: what an attention mask can isolate, and why a recurrent layer needs its own row.

LLM

RheoSampling keeps temperature sampling inside a dynamic draft tree — Proxy tree probabilities vs true verification probabilities — What does it mean?

RheoSampling gives a sampled draft token a proxy probability for the tree and its true one for verification, keeping sampling lossless.

LLM

Decompose W4A4 quantization error into correctable components — Activation-guided weight compensation vs orthogonal residual — What does it mean?

W4A4 error splits in two: the part a weight solver can absorb, and an orthogonal residual only a transform can reach.

LLM

Cactus ships Needle 3 as 2-to-20-layer subnetworks — Laddered Simple Attention Networks — What does it mean?

Needle 3 makes every layer a working model, so one 20-layer base ships as anything from 2 to 20 layers.

LLM

When2Think — Difficulty-aware Think vs NoThink routing — What does it mean?

When2Think teaches a reasoning model to decide per problem whether to think, and for how long, against that problem's own reference length.

LLM

SwitchSD reads copy intent from the model's own states — Copy-intent probe drafting-mode switch — What does it mean?

SwitchSD reads copy intent from the target model's hidden states, switching to cheap copying only when the probe predicts real repetition.

LLM

NVIDIA replaces GenAI-Perf with multiprocess AIPerf — Multiprocess load generation — What does it mean?

A single-process benchmark client is capped by Python's GIL, so the throughput you measure can be the client's ceiling, not the server's.

LLM

Let models trigger full attention only when recall helps — On-Demand Attention recall-head gating — What does it mean?

A small recall head decides, per decode step, whether reading the full KV cache is worth it — at long context, skipping it cuts decode cost.

LLM

DeepSeek-V4.1-Flash runs prefill on 8B parameters and decode on 16B — Causal Encoder-Decoder — What does it mean?

DeepSeek-V4.1-Flash runs prefill on an 8B parameter path and decode on a 16B one — a split in the parameters, not in the machines.