AI Explained

Plain explanations of trending AI concepts, with live visualizations.

Agent

MCP SEP-2106 — Full JSON Schema 2020-12 in tool I/O — What does it mean?

MCP SEP-2106 opens tool inputSchema and outputSchema to the full JSON Schema 2020-12 vocabulary — composition, conditionals, refs — and widens structuredContent from object-only to any value.

Agent

EnvFactory paper — Synthetic envs for tool-use agent training — What does it mean?

EnvFactory autonomously synthesizes 85 stateful tool environments — 5x fewer than EnvScaler / AWM — and uses topology-aware trajectory sampling to lift Qwen3 tool-use by up to 15 percentage points on BFCL v3.

LLM

Attention Once Is All You Need — Persistent KV cache across queries — What does it mean?

AOIAYN persists the KV cache across queries in a streaming session and advances it as data arrives — prefill leaves the critical path and per-query latency stays constant in context length.

LLM

ZEDA paper — Zero-output expert self-distillation — What does it mean?

ZEDA injects parameter-free zero-output experts into a finished MoE and uses two-stage self-distillation to teach the router to skip ~50% of expert FLOPs at marginal accuracy loss.

LLM

SGLang v0.5.12 — TokenSpeed MLA backend — What does it mean?

SGLang v0.5.12 ships TokenSpeed MLA — a Blackwell attention backend for Multi-head Latent Attention that caches one shared low-rank K/V latent instead of per-head K/V, with TMA bulk-store reporting up to ~12× speedup on the cache-write kernel.

Agent

MCP SEP-2468 — RFC 9207 iss parameter for OAuth mix-up defense — What does it mean?

SEP-2468 recommends MCP authorization servers include an iss parameter on auth responses, and requires clients to validate it string-equal against the recorded issuer — blocking OAuth mix-up across multi-IdP setups.

GPU

LongLive-2.0 — NVFP4 W4A4 across training and inference — What does it mean?

NVIDIA's LongLive-2.0 runs training AND inference of a 5B long-video model in NVFP4 — W4A4 matmul plus a 4-bit KV cache — for 2.15× training and 1.84× inference speedup.

LLM

Spec-decode latency paper — Load-dependent latency model — What does it mean?

Paper decomposes spec-decode latency into load-independent and load-dependent parts via Little's Law — wins shrink as the server saturates.

LLM

RoPE provably fails at long context — Position and token discrimination limits — What does it mean?

Formal proof that RoPE's attention scores converge to random along BOTH the position axis and the token-identity axis as context grows — and the RoPE base parameter only trades one collapse for the other.

Agent

RecMem paper — Subconscious + recurrence-triggered agent memory — What does it mean?

RecMem encodes every agent interaction into a cheap embedding-only 'subconscious' vector store and only invokes the LLM to consolidate clusters whose density crosses a recurrence threshold — reportedly up to 87% fewer memory-construction tokens.

Agent

MSR delegation study — Cascading fidelity loss over 20 iterations — What does it mean?

Microsoft Research's delegation stress test runs 20 rounds of LLM-to-LLM document editing with constrained in-loop verification — strong frontier models lose 19–34% artifact fidelity by iteration 20.

Agent

MCP SEP-2577 — Three deprecations and a one-year migration window — What does it mean?

SEP-2577 deprecates three early MCP features — Roots, Sampling, Logging — and introduces a Deprecated lifecycle that keeps them functional for one year, then Removed.