AI Explained
Plain explanations of trending AI concepts, with live visualizations.
COLLEAGUE.SKILL — Capability vs behavior skill tracks — What does it mean?
COLLEAGUE.SKILL turns one expert trace into a versioned skill package: a capability track (what to do) plus a behavior track (how to do it).
Agent-harness scaling law: feedback quality predicts success, not raw compute — Effective Feedback Compute (EFC) — What does it mean?
Effective Feedback Compute (EFC) predicts agent-harness success from feedback quality, not raw compute — far tighter than spend does.
Parallax — Local-linear attention vs FlashAttention 2/3 — What does it mean?
Parallax upgrades softmax attention from a flat local average to a slope-aware fit — sharper, and compute-bound past FlashAttention 2/3.
Claude Opus 4.8 — Parallel-subagent dynamic workflows — What does it mean?
Opus 4.8 'dynamic workflows' let Claude Code run parallel subagents, so wall-clock is set by the slowest subtask, not the sum.
Claude Opus 4.8 — Cache-preserving mid-task system messages — What does it mean?
Opus 4.8 can inject a system message mid-conversation without busting the prompt cache — so the cached prefix is reused, not recomputed.
OmniRetrieval — Source-native query dispatch — What does it mean?
OmniRetrieval routes each query to text, tables, or graphs natively — so JOINs and graph edges survive instead of collapsing into one flat vector index.
MarginGate — Margin-gated verification for batch-invariant decoding — What does it mean?
Why temp-0 BF16 decoding emits different tokens in a batch — and how MarginGate restores determinism by re-checking only the risky steps.
Google's Gemini Omni — Modality unification in a shared token space — What does it mean?
Gemini Omni turns text, image, audio, and video into tokens in one shared space, so a single model can read — and generate — any modality.
Parametric Memory Law links LoRA capacity to verbatim recall — The p > 0.5 recall threshold — What does it mean?
A finetune memorizes a token verbatim once its greedy probability crosses 0.5 — and a power law says how much LoRA capacity that takes.
NVIDIA AI Factories — Tokens-per-megawatt as a serving metric — What does it mean?
NVIDIA's 'AI Factories' framing reorganizes datacenter economics around tokens per megawatt — bundling compute, memory, interconnect, and orchestration into one billable knob and claiming ~50× tokens/MW for Blackwell Ultra GB300 NVL72 vs Hopper.
MobileMoE — DRAM-aware MoE scaling for sub-3GB devices — What does it mean?
MobileMoE introduces a joint memory + compute scaling law for on-device MoE LMs; S/M/L fit under 3 GB at INT4 and reportedly run 1.8–3.8× faster prefill / 2.2–3.4× faster decode than dense baselines on Galaxy S25 and iPhone 16 Pro.
Gemini 3.5 Flash — Agent-first model design — What does it mean?
Some LLMs are trained from day one to live inside an agent loop — calling tools, recovering from errors — instead of being chat models with tool-calling bolted on later.
