LLM·

LatentIndex — Shared indexer cache, per-layer token choice — What does it mean?

The news. On October 3, 2026, researchers from Peking University and Xiaohongshu posted LatentIndex (arXiv 2610.04635). It targets the indexer of DeepSeek Sparse Attention (DSA): the small scorer that, in every layer, scans the whole history to choose the 2,048 tokens the main attention will read. Training-free tests on DeepSeek-V3.2 and GLM-5 keep downstream scores close to native DSA; on DeepSeek-V3.2, four-layer sharing cuts logical indexer-cache storage by 61.1%, and a variant called hierarchical selection runs the decode indexers 2.30–2.72× faster than native DSA across 8K–128K contexts. Read the paper →

Picture a library floor with four researchers, each answering a slightly different question about the same shelf of 128,000 books. In native DSA, every researcher keeps a private card catalog of the whole shelf and reads all of it before pulling their 2,048 books. In model terms, each layer stores its own indexer key for every past token. That is the memory bill a KV cache already teaches, paid a second time by the indexer: DeepSeek-V3.2 has 61 indexer layers, each storing 132 bytes per token.

The obvious shortcut is IndexCache's: let the first researcher choose, and hand the same pile to the other three. Neighbouring layers do choose similar tokens, so this mostly works, but the followers lose the right to disagree, and the cost grows with context. At 128K tokens, the paper measures IndexCache followers losing 5.97–7.38 percentage points of attention-mass recall against what their own native indexer would have picked.

LatentIndex shares the catalog instead of the choices. The anchor layer of each four-layer group writes one compact shorthand card per token (a 192-value latent), and every layer in the group reads that same card through its own small decoder, then scores and picks its own 2,048 tokens. This is the trick MLA uses across attention heads, moved across layers. It is also a cousin of GQA, which lets several query heads share one key. Each researcher still answers their own question; they just stop copying out a private catalog first.

The catch is the cost of each lookup. Turning the shorthand back into a full key for every past token, in every layer, on every decode step would bring back the work the method was meant to remove. Instead, each layer folds its decoder into its query once per new token, so it scores the shared cards directly. This is plain linear algebra on the query-key dot product: q · (D c) equals (Dᵀ q) · c, so the decoder D (plus a small fixed bias term) is applied once to the one new query instead of to 128,000 cached latents c. In the library, the researcher translates their question into shorthand, rather than translating the whole catalog into their own words.

Sharing the cards saves memory, but each follower still reads the whole catalog. Hierarchical selection (HS) removes that scan: the anchor fills a cart with its 8,192 best candidates, and each follower picks its own 2,048 from the cart only. Followers keep their own judgment inside a smaller pool, which is why HS lands between native DSA and IndexCache on both speed and recall in the table below.

Method (DeepSeek-V3.2)What layers shareRecall (%), avg. of six lengthsIndexer time per decode token, 128KSource
Native DSANothing: own cache, own scan, own pick94.4049.60 msTable 4
IndexCacheThe token choice itself91.5413.40 msTable 4
LatentIndex (full scan)The latent cache; each layer picks93.48not reportedSec. 4.5
LatentIndex + HSLatent cache and an 8,192-token candidate cart93.1920.48 ms (2.42× vs DSA)Table 4

Four-layer sharing cuts the indexer's side cache by 61.1%, and HS cuts each follower's scoring work. Hold the model at DeepSeek-V3.2's 61 indexer layers and the context at 128K tokens (131,072). Native DSA stores 61 × 132 B = 8,052 bytes of indexer keys per token, about 1.06 GB for one sequence. LatentIndex keeps layer 0 native and turns the other 60 into 15 groups of four, each storing one 200-byte latent: 132 + 15 × 200 = 3,132 bytes per token, about 0.41 GB, which is 61.1% less. The main-attention KV cache is not touched; this is the indexer's side cache only, counted as logical bytes. On the compute side, each HS follower at 128K scores 8,192 candidates instead of 131,072 positions, 16× fewer, and the measured time for all 61 indexers per decode token (batch size one, eight GPUs) falls from 49.60 ms to 20.48 ms. IndexCache is faster still at 13.40 ms, because its followers score nothing; that 7 ms gap is the price of letting followers choose.

The trade has a second cost, measured on the full-scan version without HS. Because the anchor now reads shorthand too, even it loses a little: 1.32 points of recall at 128K, where IndexCache's anchor loses none. In exchange, average follower loss at 128K falls from 6.69 to 1.81 points, a 73% cut, and the paper finds IndexCache's loss growing with context length, which is where sparse attention matters most. The training-free version also needs a one-time calibration per model: the shared projections and decoders are fitted offline by regression and then frozen. On RULER and LongBench, two long-context benchmarks, scores stay close to native DSA's on both models tested.

Goes deeper in: LLM Internals → KV Cache → Memory Cost

Related explainers

Frequently Asked Questions

Check what you knowMap your AI & GPU knowledge across every track — free, role-based