iS-KV — Basis-synchronized online low-rank KV compression — What does it mean?
The news. On October 2, 2026, researchers from HKUST (Guangzhou), the Shenzhen Institutes of Advanced Technology, the University of Macau and OPPO posted iS-KV: Online Low-Rank KV Cache Compression via Block-Incremental SVD on arXiv. They test it on DeepSeek-R1-Distill-Llama-8B and Qwen3-8B, mainly on the MATH-500 math benchmark with a 16K-token limit, and compare it with the eviction methods R-KV and SnapKV at matched memory. At all six matched-memory settings, iS-KV scored higher than both eviction baselines. Read the paper →
Picture a paint shop that must be able to remake every color it has ever mixed, but has counter space for only a few real swatches. One option is to throw old swatches away. That is token eviction, and a customer who asks for a thrown-away color is out of luck. The other option is to keep a small set of base paints and write each old color down as a recipe: so much of paint one, so much of paint two, so much of paint three. iS-KV takes the second option for the KV cache: every token position stays, but older keys and values are stored as short coefficient vectors over a shared basis instead of as full vectors. The newest tokens, a recent window of 128 in the paper's main runs, stay exact like the swatches on the counter, and a set of protected prompt positions also stays exact.
Why not just bin the quiet colors? In long reasoning, a token the model ignores now is often one it needs later. The authors marked the 75% of tokens with the lowest recent attention at decoding step 1,024, removed nothing, and kept watching: on the AIME math problems about 30% of those tokens received strong attention again within the next 8K tokens, and about 60% did by 30K tokens. Every new query scores every earlier key, so a deleted key is something the model can no longer look at. Keeping old tokens in compressed form is possible because generated keys and values are strongly low-rank: a handful of directions explain most of their variation, and keys show this most clearly before RoPE rotates them by position. The diagram below shows the eviction policies that iS-KV replaces.
Remove earliest tokens first
Now the hard part of the paint shop. New colors keep arriving: every time 64 tokens leave the recent window, they are folded into the compressed history as one block. The best base paints for the old colors are not the best base paints for old plus new, so the basis is updated. If you change the base paints but keep the old recipes, every old recipe now mixes a different color, and the stored history drifts even though nothing touched it. The paper measures this. At step 256, updating only the basis gave a drift of 1.55 for keys and 1.40 for values on the paper's drift measure; updating the basis and the old coefficients together cut it to 0.023 and 0.051. Drift here is the change in the stored history caused by one update, not the error against the original uncompressed cache. iS-KV does the joint update with block-incremental SVD, a classical method by Brand (2006) that rewrites the basis and every stored coefficient in one step. The decomposition itself is solved on a small core matrix whose size does not depend on the history length, although rewriting the stored coefficients still costs more as the history grows.
A worked example from the paper, holding three things fixed: Qwen3-8B, the MATH-500 benchmark, and a persistent KV budget of about 17.7% of the full cache (5.64× compression), counted in bytes the same way for every method. Uncompressed, the model answers 94.4% correctly. R-KV, which keeps a selected subset of tokens inside that budget, drops to 68.8%. iS-KV at rank 96, which keeps every position in exact or low-rank form, reaches 89.2%. Same memory budget, and keeping every token approximately beats keeping some tokens exactly by 20.4 points. The matching counts the low-rank factors, the exact rows and R-KV's 128-token staging buffer, and excludes model weights and temporary rebuild buffers.
| Model (MATH-500, matched memory) | Compression | Full cache | R-KV | SnapKV | iS-KV | Source |
|---|---|---|---|---|---|---|
| DeepSeek-R1-Distill-Llama-8B | 4.06× | 83.6% | 78.2% | 70.1% | 82.6% (rank 128) | Table 1 |
| DeepSeek-R1-Distill-Llama-8B | 5.92× | 83.6% | 72.6% | 64.3% | 78.2% | Table 1 |
| Qwen3-8B | 5.64× | 94.4% | 68.8% | 73.17% (reported) | 89.2% (rank 96) | Table 1 |
| Qwen3-8B | ≈7.1× | 94.4% | 60.8% | 65.04% (reported) | 87.8% (rank 64) | Table 1 |
The trade is that memory still grows and generation gets slower. Each compressed token is cheaper, but each one is still stored, so iS-KV slows cache growth rather than capping it. In a single-run timing test on one NVIDIA A40 (DeepSeek-R1-Distill-Llama-8B, rank 128, 1,024 decode steps after the input), KV memory fell from 4.125 GiB to 1.808 GiB with a 32K input, a 56.2% saving that shrinks to 33.7% with an 8K input, while total generation time rose by 8.8% to 22.2% across the 8K to 32K inputs. The authors call these single runs rather than a stable latency measurement, and they did not test low rank combined with low-bit quantization or inside a serving engine. Low rank is a third lever next to two the curriculum already covers: GQA cuts the number of KV heads, quantization cuts the bits per number, and iS-KV cuts the number of coordinates stored per old token.
Goes deeper in: LLM Internals → KV Cache → Memory Cost
Related explainers
- KV-COBRA: rank vs bit-width allocation per head — a fixed low-rank projection whose rank is chosen per head; iS-KV instead moves its basis online during decoding.
- Risk-controlled KV cache eviction — the eviction side of the trade: how much cache to drop, certified against a failure target.
- WorldKV: evict-and-reinsert KV memory — another answer to "old tokens come back", which stores evicted chunks and reinserts them.