AI Explained

Plain explanations of trending AI concepts, with live visualizations.

Agent

NemotronLabs VoiceChat listens while it speaks — Full-duplex vs half-duplex turn-taking — What does it mean?

A full-duplex speech model listens and speaks on parallel channels, so silence becomes a trained output rather than a gap between turns.

Agent

ActObs unmasks the environment's half of an agent trajectory — Observation-token supervision — What does it mean?

Agent SFT grades only the agent's commands. ActObs also supervises the environment's replies - about 45% of the trajectory tokens.

LLM

dQwen3.5 — Adapting a hybrid AR backbone into a diffusion LM — What does it mean?

Only a quarter of Qwen3.5's layers are attention. Un-mask just those and the hybrid becomes a diffusion LM in about half the tokens.

Agent

Agent skill memory that rewrites itself from user traffic — Matched replay gate — What does it mean?

An agent rewrites its own skill library, and a matched replay gate refuses any edit that regresses a single replayed case.

LLM

Kev reads a document once and answers every question about it separately — Question-isolation masks — What does it mean?

Read the document once, answer every question separately: what an attention mask can isolate, and why a recurrent layer needs its own row.

LLM

RheoSampling keeps temperature sampling inside a dynamic draft tree — Proxy tree probabilities vs true verification probabilities — What does it mean?

RheoSampling gives a sampled draft token a proxy probability for the tree and its true one for verification, keeping sampling lossless.

LLM

Decompose W4A4 quantization error into correctable components — Activation-guided weight compensation vs orthogonal residual — What does it mean?

W4A4 error splits in two: the part a weight solver can absorb, and an orthogonal residual only a transform can reach.

LLM

Cactus ships Needle 3 as 2-to-20-layer subnetworks — Laddered Simple Attention Networks — What does it mean?

Needle 3 makes every layer a working model, so one 20-layer base ships as anything from 2 to 20 layers.

Agent

A 176-setting study ablates the coding-agent harness — Harness component ablation — What does it mean?

Benchmarking a whole harness hides which part earned the gain. Ablate one component and the answer changes with model and budget.

LLM

When2Think — Difficulty-aware Think vs NoThink routing — What does it mean?

When2Think teaches a reasoning model to decide per problem whether to think, and for how long, against that problem's own reference length.

Agent

LangChain benchmarks Jev — Evaluator variance vs the generative LLM judge — What does it mean?

A grader has error of its own. Freeze the agent's trace, replay the grader, and the spread you see is entirely the grader's.

Agent

Schedule agent tools by task-specific CPU bottlenecks — Task-aware CPU admission vs uniform resource allocation — What does it mean?

An agent's bottleneck is often local CPU or disk, not the model — so size each task's container by what its tools actually need.