AI Explained
Plain explanations of trending AI concepts, with live visualizations.
D-Quant drifts variable-length KV codes into fixed-size token streams — Entropy-coded KV cache — What does it mean?
Short codewords for common KV values, long for rare ones, then drifted back into fixed-size per-token streams.
JustFit serves a 212,992-token context on a 24 GiB laptop — Just-in-time state residency — What does it mean?
JustFit holds a 212,992-token context on a 24 GiB laptop by scheduling when each piece of state is resident, not by shrinking the weights.
FlexEE cuts LLM decoding by up to 3.16x under weight offloading — KV-compatible early exit — What does it mean?
Stopping a token at layer 12 saves compute and a weight fetch — but only if the layers it skipped still leave a usable KV cache.
Cartridges match in-context learning once retrieval is real — KV cartridges vs parametric fine-tuning — What does it mean?
In multi-document retrieval, a trained KV cartridge is the only method that matches leaving the documents in the context window.
LoopSpec drafts the next token while it verifies this one — Pipelined self-speculative decoding — What does it mean?
A looped Transformer can read a draft off its own early passes, and LoopSpec makes that draft happen while the full depth verifies.
MAPS cuts LLM tail latency up to 84.8% — Uncertainty-calibrated output-length bounds — What does it mean?
Output length is unknown when a request arrives. MAPS hands the scheduler a ceiling with a chosen error rate instead.
Stream an 8B MoE from SSD in 1 GiB active memory — One-step-ahead expert prerouting — What does it mean?
Predict the next token's experts one step early so the SSD read hides under the current forward pass.
Qwen3.8-Flash-Next widens the residual stream into four gated branches — Gated Residual — What does it mean?
Qwen splits the transformer's shared residual stream into four branches and gates which ones each layer reads and writes.
Qwen3.8-Flash-Next adds 51B of embeddings that can live off the GPU — N-gram embedding offload — What does it mean?
Qwen puts 51B parameters in an n-gram-keyed lookup table and moves it off the GPU, prefetching rows over PCIe.
SQD splits decode by attention type, not by operator — Attention-type decode partitioning — What does it mean?
SQD cuts decode by attention complexity, so full-KV scanning and fixed-footprint math land on machines sized for each.
LILA prunes LLM neurons without calibration data — Calibration-free structured neuron pruning — What does it mean?
LILA ranks FFN neurons for deletion from the weight matrix alone: a closed-form spectral score, no calibration data, no gradients.
LOCUS cuts LLM output length by up to 39.84% — Utility-constrained length reduction — What does it mean?
LOCUS runs the same preference objective inside a constrained low-rank subspace; answers came out up to 39.84% shorter.