AI Explained

Plain explanations of trending AI concepts, with live visualizations.

GPU

NVIDIA Jetson Thor — Edge Blackwell vs datacenter Blackwell — What does it mean?

Jetson Thor is NVIDIA's Blackwell-architecture edge AI module — 2,070 FP4 TFLOPS in a 40–130W envelope, reportedly 7.5× compute and 3.5× per-watt vs Jetson Orin (Computex 2026).

Agent

Anthropic's Project Glasswing — Detection-saturated vulnerability pipeline — What does it mean?

Glasswing partners found 10,000+ high- or critical-severity vulnerabilities in one month. The detection-saturated dynamic is sharpest where the patcher is a separate org — Anthropic's OSS disclosures: 530 reported, only 75 patched after a month.

Agent

PromptArmor × Copilot Cowork — Image-URL exfiltration in agent UIs — What does it mean?

PromptArmor showed Microsoft Copilot Cowork posting a Teams message whose hidden image tag leaks a pre-authenticated OneDrive token the moment the user opens the inbox — zero clicks on a malicious link.

LLM

ThriftAttention paper — Importance-aware FP16/FP4 mixed-precision attention — What does it mean?

ThriftAttention runs ~5% of QK attention blocks (the top by a cheap importance score) in FP16 and the rest in FP4, recovering 89.1% of the FP4→FP16 long-context quality gap.

Agent

PushBench paper — Quantitative Goal Persistence (QGP) — What does it mean?

PushBench measures Quantitative Goal Persistence — frontier models hold up at 50 artifacts but drop to 3/9 successes at 100; a state-tracking controller restores 69–78% by rejecting duplicate submissions.

LLM

Distillation in LLM pre-training — Non-monotonic teacher strength — What does it mean?

Sweeping teacher and student sizes in pre-training distillation shows the helpfulness of a teacher peaks at small undertrained models — stronger teachers can hurt, and the gains land out-of-domain, not in-domain.

GPU

I/O-optimal approximate attention — Near-linear I/O vs FlashAttention — What does it mean?

A new paper derives approximate-attention algorithms whose I/O between SRAM and HBM scales near-linearly in sequence length n — vs FlashAttention's quadratic n² — with matching I/O lower bounds proving the result is near-optimal.

LLM

Complete-muE paper — Two-bridge muTransfer for MoE — What does it mean?

Complete-muE extends muTransfer to MoE via two bridges — active-width scaling (dense → activated width) and activated-expert scaling (across MoE shapes) — so one dense sweep transfers near-optimally to any expert count and top-k.

LLM

Cursor Composer 2.5 — Targeted textual feedback RL — What does it mean?

Cursor Composer 2.5 ships targeted textual feedback RL — a constructed short hint at a specific span in a long agent rollout becomes a teacher distribution, and on-policy distillation KL replaces the diffuse end-of-rollout scalar with a localized credit signal.

LLM

VPO paper — Vector-reward advantage vs GRPO scalar collapse — What does it mean?

VPO replaces GRPO's scalar advantage estimator with a vector-valued one — so the post-trained policy keeps a diverse solution distribution that pays off at pass@k and best@k as the search budget grows.

Agent

OpenSCAD Pantheon benchmark — Human-in-the-loop vs autonomous coding agents — What does it mean?

ModelRift pitted 6 agentic coding tools against the same OpenSCAD Pantheon task — Antigravity 2.0 in autonomous mode won at 4.5/5 in ~12 min while ModelRift's human-in-the-loop tier hit 3.8/5 in ~10 min.

Agent

MCP 2026-07-28 RC — stateless transport — What does it mean?

MCP's 2026-07-28 RC reworks transport so every tools/call request carries its own routing data — any server in the fleet can serve it, no sticky session pin required.