Plain explanations of trending AI concepts, with live visualizations.
NVIDIA's LongLive-2.0 runs training AND inference of a 5B long-video model in NVFP4 — W4A4 matmul plus a 4-bit KV cache — for 2.15× training and 1.84× inference speedup.