AI Explained

Plain explanations of trending AI concepts, with live visualizations.

LLM

D-Quant drifts variable-length KV codes into fixed-size token streams — Entropy-coded KV cache — What does it mean?

Short codewords for common KV values, long for rare ones, then drifted back into fixed-size per-token streams.

LLM

JustFit serves a 212,992-token context on a 24 GiB laptop — Just-in-time state residency — What does it mean?

JustFit holds a 212,992-token context on a 24 GiB laptop by scheduling when each piece of state is resident, not by shrinking the weights.

LLM

FlexEE cuts LLM decoding by up to 3.16x under weight offloading — KV-compatible early exit — What does it mean?

Stopping a token at layer 12 saves compute and a weight fetch — but only if the layers it skipped still leave a usable KV cache.

LLM

Cartridges match in-context learning once retrieval is real — KV cartridges vs parametric fine-tuning — What does it mean?

In multi-document retrieval, a trained KV cartridge is the only method that matches leaving the documents in the context window.

LLM

LoopSpec drafts the next token while it verifies this one — Pipelined self-speculative decoding — What does it mean?

A looped Transformer can read a draft off its own early passes, and LoopSpec makes that draft happen while the full depth verifies.

LLM

MAPS cuts LLM tail latency up to 84.8% — Uncertainty-calibrated output-length bounds — What does it mean?

Output length is unknown when a request arrives. MAPS hands the scheduler a ceiling with a chosen error rate instead.

LLM

Stream an 8B MoE from SSD in 1 GiB active memory — One-step-ahead expert prerouting — What does it mean?

Predict the next token's experts one step early so the SSD read hides under the current forward pass.

LLM

Qwen3.8-Flash-Next widens the residual stream into four gated branches — Gated Residual — What does it mean?

Qwen splits the transformer's shared residual stream into four branches and gates which ones each layer reads and writes.

LLM

Qwen3.8-Flash-Next adds 51B of embeddings that can live off the GPU — N-gram embedding offload — What does it mean?

Qwen puts 51B parameters in an n-gram-keyed lookup table and moves it off the GPU, prefetching rows over PCIe.

LLM

SQD splits decode by attention type, not by operator — Attention-type decode partitioning — What does it mean?

SQD cuts decode by attention complexity, so full-KV scanning and fixed-footprint math land on machines sized for each.

LLM

LILA prunes LLM neurons without calibration data — Calibration-free structured neuron pruning — What does it mean?

LILA ranks FFN neurons for deletion from the weight matrix alone: a closed-form spectral score, no calibration data, no gradients.

LLM

LOCUS cuts LLM output length by up to 39.84% — Utility-constrained length reduction — What does it mean?

LOCUS runs the same preference objective inside a constrained low-rank subspace; answers came out up to 39.84% shorter.