AI Explained
Plain explanations of trending AI concepts, with live visualizations.
TELBench localizes where deep-research agents go wrong — Span-level error localization — What does it mean?
TELBench + DRIFT pinpoint which step of a deep-research agent's trajectory made the answer unreliable — span-level error localization, up to +30 pp.
StreamMA — Streaming inter-agent reasoning — What does it mean?
StreamMA streams each reasoning step between agents — B starts on A's reliable early steps, not the risky last one. +7.3pp avg accuracy.
NVIDIA RTX Spark superchip — Unified CPU–GPU memory — What does it mean?
RTX Spark wires a Grace CPU and a Blackwell GPU to one 128GB pool over NVLink-C2C, so the GPU skips the PCIe host–device copy.
Microsoft MAI-Code-1-Flash — Adaptive solution-length control — What does it mean?
MAI-Code-1-Flash scales how many reasoning tokens it spends to each task's difficulty — short chains for easy tasks, long only for hard ones.
KVarN squeezes the KV cache to 2 bits — Hadamard rotation — What does it mean?
KVarN rotates outliers out of the KV cache so 2-bit quantization fits every channel — calibration-free, no error compounding across decode.
Crafter paper — Multi-agent refinement harness with a directive critic — What does it mean?
Crafter's directive critic emits per-dimension fixes + typed edits, not a scalar score — a harness that lifts figures 33.73 → 50.34.
WASH attack washes out LLM text watermarks — Watermark removal by model-averaging — What does it mean?
Averaging the output distributions of 3–5 independent LLMs cancels each one's text watermark — detection z-scores fall from 5–300 to below 2.
PEFT scaling paper — Persistent personal adapters at million-scale — What does it mean?
Reframes a LoRA adapter from a cost-cutting trick into persistent per-user state — a million personal adapters served over one frozen base.
LongTraceRL — Rubric reward (entity-level process supervision) — What does it mean?
LongTraceRL scores each reasoning hop, not just the final answer — dense process reward, gated to correct rollouts so it cannot be gamed.
Harness-1 — State-externalizing search harness — What does it mean?
Harness-1 is a 20B search agent that keeps working memory in an external harness, not a growing transcript — so context stays flat as the search deepens.
GrepSeek trains a search agent to use shell commands — GRPO-trained shell-command search — What does it mean?
GrepSeek trains an agent to search a raw corpus with shell commands via a Tutor/Planner distillation then GRPO — index-free agentic retrieval.
dMoE cuts diffusion-LLM MoE memory ~80% — block-level expert routing — What does it mean?
dMoE pools a diffusion block's per-token expert choices into one block-level decision — ~70→15 experts loaded, ~80% less memory.