AI Explained
Plain explanations of trending AI concepts, with live visualizations.
INT8 finally beats FP8 on consumer GPUs — Fused INT8 GEMM kernel — What does it mean?
A fused Triton kernel keeps INT8 matmuls on the tensor cores end to end, so W8A8 finally beats FP8 on a consumer GPU — no dequant round trip.
CacheRL trains tool-calling agents via cached rollouts at 100× less compute — Cached rollouts for agent RL — What does it mean?
CacheRL replaces live tool execution during RL rollouts with a three-tier fuzzy cache — 92% process accuracy vs GPT-5's 94% at ~100× less compute.
HarnessBridge — Learned agent harness vs hand-engineered — What does it mean?
HarnessBridge replaces the hand-built agent harness with a learnable module — two projections that distill state and vet each action.
EvoArena + EvoMem — Patch-based agent memory — What does it mean?
EvoMem keeps agent memory as a changelog of structured patches — what changed and when — so the agent can reason about how its world evolved.
NVIDIA Blackwell leads AgentPerf, the first agentic-AI infra benchmark — Trajectory-replay benchmarking — What does it mean?
AgentPerf grades serving systems by replaying real multi-step agent runs, not single prompts — and Blackwell's GB300 leads on agents per megawatt.
WeaveBench: best computer-use agent clears just 41% — Trajectory-aware vs outcome-only grading — What does it mean?
WeaveBench finds the best computer-use agent clears 41.2% — and a trajectory-aware judge shows outcome-only grading flatters the rest.
VIA-SD speeds up speculative decoding 10–20% — Tiered confidence-gated verification — What does it mean?
VIA-SD adds a confidence-gated middle tier to speculative decoding — close calls go to a slim sub-network of the same model, not a full re-run.
SpatialClaw lifts agent spatial reasoning to 59.9% — Code-as-action vs structured tool-calls — What does it mean?
SpatialClaw makes a VLM agent's actions executable Python cells on a stateful kernel — observe-then-act beats rigid tool-calls, +11.2 pts to 59.9%.
MaxProof clears IMO/USAMO gold — Defense-in-depth generative verifier — What does it mean?
MaxProof tunes its proof verifier for a very low false-positive rate, so sampling many candidate proofs and picking a winner by tournament actually works.
A survey of agent-environment engineering — Symbolic vs neural environment synthesis — What does it mean?
A survey reframes building an agent's training world as engineering — and its sharpest split is hand-coded vs model-generated environments.
Manifold Power Iteration redesigns MoE routers — Router-to-expert alignment — What does it mean?
Manifold Power Iteration rotates each MoE router row onto its expert's top singular direction — better routing at 0.2% train cost, zero inference overhead.
CodeSpear strips an LLM's ability to refuse — Grammar-constrained decoding jailbreak — What does it mean?
Force an LLM's output to fit a code grammar and its natural-language refusal becomes invalid — CodeSpear uses this to lift attack success to ~82%.











