AI Explained
Plain explanations of trending AI concepts, with live visualizations.
Cut ASR GPU infrastructure 75% with CUDA MPS — Concurrent execution vs time-slicing — What does it mean?
CUDA MPS lets several processes share one GPU at once instead of taking turns, turning idle SMs into throughput.
MeanField surrogate schedules concurrent models on a shared GPU — Mean-field contention surrogate — What does it mean?
Predicting how colocated models slow each other down carries a combinatorial profiling cost. MeanField reads aggregate GPU state instead.
NVIDIA sizes speculative-decoding drafts to the GPU's attention tile — Tile-aligned draft length — What does it mean?
When attention dominates the step, NVIDIA picks draft length D = 128 / G - 1 so the verification pass exactly fills a 128-row GPU tile.
Protect LLM training from silent corruption with 1.65–6.76% overhead — Silent data corruption in transformer training — What does it mean?
A silent bit flip does not crash training. TrainSDC guards only the interfaces that amplify a fault, at 1.65-6.76% runtime overhead.
NVIDIA pairs Rubin GPUs with Groq 3 LPX — Phase-split inference across two accelerator types — What does it mean?
Prefill is compute-bound and decode is memory-bandwidth-bound. NVIDIA now runs each phase on a different kind of chip inside one rack.
Microsoft details the Maia 200 inference accelerator — Software-defined dataflow — What does it mean?
A GPU decides at runtime what runs next. Maia 200 moves that into the compiler, which writes the data-movement schedule ahead of time.
Target 10,000 tok/s with Cerebras CS-5 and 3D-stacked memory in CS-6 — Wafer-scale memory locality — What does it mean?
Decode re-reads the weights for every token. Cerebras' answer is to stop moving them: keep the weights on the same wafer as the math.
TokenStack moves hot KV blocks into HBM-PIM — Processing-in-memory attention — What does it mean?
Most KV-cache work shrinks the bytes so fewer of them travel. TokenStack instead puts attention hardware on the memory layers the hot blocks already sit on.
RMM cuts Transformer matmuls without touching the weights — Contraction-dimension slicing — What does it mean?
Most inference savings shrink the numbers or skip whole layers. RMM keeps the weights and drops slices of the dimension a matmul sums over.
Intel reveals a 480GB air-cooled inference GPU — Capacity-first inference memory — What does it mean?
Almost every inference accelerator buys bandwidth with HBM. Crescent Island reaches for capacity instead, with 480GB of slower LPDDR5X.
AsmEvo tunes AMD GPU kernels at the assembly level — Correctness-gated assembly optimization — What does it mean?
AsmEvo lets an agent rewrite AMDGPU assembly, keeping only the edits whose output still matches the original on identical launches.
Atrex-Bench tests LLM-written kernels on production traces — Trace-weighted kernel benchmarking — What does it mean?
A kernel benchmark that samples operators and shapes from real serving traces and weights each one by the GPU time it actually burns.







