AI Explained
Plain explanations of trending AI concepts, with live visualizations.
vLLM 0.30 ships Fast Start — Persistent GPU weight cache via CUDA IPC — What does it mean?
vLLM 0.30 keeps prepared weights in a per-GPU daemon, so a restarted engine maps them over CUDA IPC instead of reloading them from disk.
Qwen-Image-2.1 runs two mask rules in one sequence — Mixed-granularity attention — What does it mean?
One sequence, two mask granularities: causal for text, chunk-level for image — plus a static-context KV cache computed once.
dQwen3.5 — Adapting a hybrid AR backbone into a diffusion LM — What does it mean?
Only a quarter of Qwen3.5's layers are attention. Un-mask just those and the hybrid becomes a diffusion LM in about half the tokens.
Kev reads a document once and answers every question about it separately — Question-isolation masks — What does it mean?
Read the document once, answer every question separately: what an attention mask can isolate, and why a recurrent layer needs its own row.
RheoSampling keeps temperature sampling inside a dynamic draft tree — Proxy tree probabilities vs true verification probabilities — What does it mean?
RheoSampling gives a sampled draft token a proxy probability for the tree and its true one for verification, keeping sampling lossless.
Decompose W4A4 quantization error into correctable components — Activation-guided weight compensation vs orthogonal residual — What does it mean?
W4A4 error splits in two: the part a weight solver can absorb, and an orthogonal residual only a transform can reach.
Cactus ships Needle 3 as 2-to-20-layer subnetworks — Laddered Simple Attention Networks — What does it mean?
Needle 3 makes every layer a working model, so one 20-layer base ships as anything from 2 to 20 layers.
When2Think — Difficulty-aware Think vs NoThink routing — What does it mean?
When2Think teaches a reasoning model to decide per problem whether to think, and for how long, against that problem's own reference length.
SwitchSD reads copy intent from the model's own states — Copy-intent probe drafting-mode switch — What does it mean?
SwitchSD reads copy intent from the target model's hidden states, switching to cheap copying only when the probe predicts real repetition.
NVIDIA replaces GenAI-Perf with multiprocess AIPerf — Multiprocess load generation — What does it mean?
A single-process benchmark client is capped by Python's GIL, so the throughput you measure can be the client's ceiling, not the server's.
Let models trigger full attention only when recall helps — On-Demand Attention recall-head gating — What does it mean?
A small recall head decides, per decode step, whether reading the full KV cache is worth it — at long context, skipping it cuts decode cost.
DeepSeek-V4.1-Flash runs prefill on 8B parameters and decode on 16B — Causal Encoder-Decoder — What does it mean?
DeepSeek-V4.1-Flash runs prefill on an 8B parameter path and decode on a 16B one — a split in the parameters, not in the machines.