AI Explained

Plain explanations of trending AI concepts, with live visualizations.

GPU

AMDKernelVault trains an 8B model to write AMD GPU kernels — Hierarchical execution reward — What does it mean?

Binary pass/fail leaves most GRPO rounds with zero advantage. AMDKernelVault grades a kernel 0.4, 1.0 and up to 1.5 instead.

GPU

Devin-built GPU sieving factored RSA-260 — Preemptible sieving vs the interconnect-bound solve — What does it mean?

Sieving survives on preemptible GPUs because losing a unit costs milliseconds. The block Wiedemann solve cannot, so only it pays for NVLink.

GPU

NVLink 6 multi-layer resiliency — Recovery escalation ladder — What does it mean?

NVLink 6 fixes a fault on the cheapest rung that can: FEC under 1 ms, software 1.5 s, shadow engine 7.3 s, cold restart 283 s.

GPU

OPEN-1B publishes a training run you can replay bit for bit — Step-replay auditing — What does it mean?

OPEN-1B orders every reduction, batch and collective so one training step — not the whole run — becomes the unit anyone can verify.

GPU

Decoupling prefill and decode power saves 32.3% of a lane pair's electricity — Phase-decoupled power control — What does it mean?

Prefill and decode sit on opposite sides of the roofline, so one power setting for both is wrong for one. Give each lane its own rule.

GPU

Replace exp-then-quantize softmax to cut vector latency 40.33% — Exp-free softmax quantization to E2M1 probability codes — What does it mean?

EFQ-Softmax emits 4-bit E2M1 probabilities directly from attention scores, deleting the exp() stage inside FlashAttention.

GPU

TPUv7 Ironwood serving benchmark — PrivateUse1 backend with retained Pallas kernels — What does it mean?

PyTorch reaches the TPU through a reserved dispatch slot, compiles via StableHLO, and retains hand-written kernels.

GPU

TPUv7 Ironwood serving benchmark — Data-parallel attention vs expert parallelism — What does it mean?

Two KV heads leave nothing to split, so attention goes data-parallel while 512 experts stay expert-parallel.

GPU

Fuse decode into one megakernel for 1.58× higher H100 throughput — Wave quantization — What does it mean?

When a kernel's tile count is not a multiple of the SM count, the last wave runs half-empty and you pay for a full wave anyway.

GPU

Fuse decode into one megakernel for 1.58× higher H100 throughput — Persistent decode megakernel — What does it mean?

One persistent CUDA kernel runs a whole decode step, deleting the kernel boundaries instead of batching their launches.

GPU

NVIDIA opens two Rust paths for CUDA kernels — SIMT thread model vs tile-level programming — What does it mean?

NVIDIA opened two Rust paths to CUDA kernels, one per-thread and one per-tile. The split contrasts two established GPU programming models.

GPU

PyTorch 2.14 ships NVGEMM kernels — NVGEMM epilogue fusion — What does it mean?

PyTorch 2.14 adds a third matmul backend that fuses the bias and activation into the GEMM, and lets the compiler measure which kernel wins.