AI Explained
Plain explanations of trending AI concepts, with live visualizations.
Agentic Botnets exploits hallucinated repo and skill names — Adversarial hallucination squatting — What does it mean?
Attackers pre-register the fake repo and skill names an LLM hallucinates, then hide malicious prompts there — up to 85% repo / 100% skill hallucination rates.
ToolFailBench separates skip, ignore, and fabricate failures — Tool-use failure taxonomy — What does it mean?
ToolFailBench sorts agent tool-use failures into four modes, and adds trap tasks to catch tools that shouldn't have been used.
Vera tests LLM-agent safety with executable evidence checks — Evidence-grounded verification — What does it mean?
Vera judges an AI agent's safety from the real evidence it leaves in a sandbox — files, tool calls, commands — not the model's own say-so.
SIS turns off-policy RL tokens into on-policy updates — Selective Importance Sampling — What does it mean?
Selective Importance Sampling accepts the agreeing tokens as on-policy (ratio 1), taming the variance when RL reuses off-policy rollouts.
LLM-as-a-Verifier scales agent feedback with logit-score expectations — Verification as a scaling axis — What does it mean?
Reading the verifier's whole score distribution — not one label — turns a coarse pass/fail into a continuous score you can scale.
Discrete diffusion theory unifies denoisers, scores & bridge predictors — One object, three coordinates — What does it mean?
The three ways to train a diffusion language model — denoiser, score, and bridge — turn out to be one object written in different coordinates.
Direct-OPD transfers weak-model RL gains as log-ratio rewards — Weak-to-strong reward transfer — What does it mean?
Direct-OPD reuses a small model's before/after RL log-ratio as a dense reward to lift a stronger model — no RL on the big model.
Omnigent open-sources a meta-harness for coding agents — The agent meta-harness — What does it mean?
A meta-harness is one control layer above many coding-agent CLIs — it holds credentials, session, policy, and sandbox in one place.
LOCOS finds non-literal retrieval heads by scoring logit contribution — Logit-Contribution Scoring — What does it mean?
LOCOS finds the attention heads that answer from meaning, not copied words, by scoring how much each head writes toward the answer token.
LangChain adds dynamic subagents for code-driven orchestration — Programmatic subagent fan-out — What does it mean?
LangChain Deep Agents can now write a short script that fans out one subagent per chunk, so coverage is a property of code, not a prompt.
LACUNA tests whether LLM unlearning hits the right weights — Output-level vs weight-level unlearning evaluation — What does it mean?
LACUNA plants a fact in known weights to test whether unlearning erases it, or just hides the output while a resurfacing attack revives it.
HaloGuard ships 0.8B open constitutional safety classifier — Paired counterfactual safety data — What does it mean?
HaloGuard trains a tiny safety classifier on matched prompt pairs that flip only intent, so it learns the boundary, not the keywords.











