AI Explained
Plain explanations of trending AI concepts, with live visualizations.
NemotronLabs VoiceChat listens while it speaks — Full-duplex vs half-duplex turn-taking — What does it mean?
A full-duplex speech model listens and speaks on parallel channels, so silence becomes a trained output rather than a gap between turns.
ActObs unmasks the environment's half of an agent trajectory — Observation-token supervision — What does it mean?
Agent SFT grades only the agent's commands. ActObs also supervises the environment's replies - about 45% of the trajectory tokens.
dQwen3.5 — Adapting a hybrid AR backbone into a diffusion LM — What does it mean?
Only a quarter of Qwen3.5's layers are attention. Un-mask just those and the hybrid becomes a diffusion LM in about half the tokens.
Agent skill memory that rewrites itself from user traffic — Matched replay gate — What does it mean?
An agent rewrites its own skill library, and a matched replay gate refuses any edit that regresses a single replayed case.
Kev reads a document once and answers every question about it separately — Question-isolation masks — What does it mean?
Read the document once, answer every question separately: what an attention mask can isolate, and why a recurrent layer needs its own row.
RheoSampling keeps temperature sampling inside a dynamic draft tree — Proxy tree probabilities vs true verification probabilities — What does it mean?
RheoSampling gives a sampled draft token a proxy probability for the tree and its true one for verification, keeping sampling lossless.
Decompose W4A4 quantization error into correctable components — Activation-guided weight compensation vs orthogonal residual — What does it mean?
W4A4 error splits in two: the part a weight solver can absorb, and an orthogonal residual only a transform can reach.
Cactus ships Needle 3 as 2-to-20-layer subnetworks — Laddered Simple Attention Networks — What does it mean?
Needle 3 makes every layer a working model, so one 20-layer base ships as anything from 2 to 20 layers.
A 176-setting study ablates the coding-agent harness — Harness component ablation — What does it mean?
Benchmarking a whole harness hides which part earned the gain. Ablate one component and the answer changes with model and budget.
When2Think — Difficulty-aware Think vs NoThink routing — What does it mean?
When2Think teaches a reasoning model to decide per problem whether to think, and for how long, against that problem's own reference length.
LangChain benchmarks Jev — Evaluator variance vs the generative LLM judge — What does it mean?
A grader has error of its own. Freeze the agent's trace, replay the grader, and the spread you see is entirely the grader's.
Schedule agent tools by task-specific CPU bottlenecks — Task-aware CPU admission vs uniform resource allocation — What does it mean?
An agent's bottleneck is often local CPU or disk, not the model — so size each task's container by what its tools actually need.