CoWindow Attention splits distant history across heads — Complementary per-head windows — What does it mean?
The news. On September 26, 2026, researchers from HKUST (Guangzhou), the Beijing Academy of Artificial Intelligence and Université Paris Cité posted CoWindow Attention: Full Causal Coverage Is a Collective Property (arXiv 2609.32704). They train models from 0.6B to 14B parameters with the pattern, plus 32B models through continued training, and report quality comparable to full attention. In an attention-operator benchmark at 128K tokens on 8 H100 GPUs with tensor parallelism, they report 7.4× faster training forward, 8.6× faster backward and 3.0× faster decoding than full attention. The code is released as flash-sparse-attention. Read the paper →
Picture a long beach with eight lifeguards. The naive rota tells every guard to scan the whole beach, from the flag at the far end to the water in front of their own tower. That is what standard multi-head attention does: every head computes a score against every earlier token, so at 128K tokens each head re-reads 128K keys and the eight heads read the same far water eight times.
The cheap fix, sliding-window attention, tells every guard to watch only the water near the tower. That cuts the work, but now nobody watches the far end at all. In attention terms, the causal mask is narrowed to the last W tokens for every head, and a fact stated 100K tokens ago cannot be read directly by any head in that layer.
CoWA keeps the shared near water and the start flag for every guard, then cuts the far water into eight stretches and gives each guard exactly one. No guard sees the whole beach, yet every stretch has a guard on it. The paper calls this collective coverage: for each query, the union of what the heads can see is the full causal prefix, even though each head is sparse over the distant history.
The rule is set by position, not learned. A head's windows are fixed by the current sequence length and its global KV-head index, so the kernel knows which key blocks to visit before computing any score. That is the difference from routed sparse attention such as MoBA or DeepSeek's DSA, which first score blocks or tokens to decide what to keep. CoWA has no router or indexer to pay for: excluded blocks are skipped before the QK product, the same block-skipping idea a fused FlashAttention kernel already uses for the causal triangle.
The position rule also fits how large models are split across GPUs. Under tensor parallelism, each GPU holds a contiguous group of KV heads. Because windows are assigned by the global head index, each GPU owns a different slice of the far past; only the small shared near and sink windows are read on every GPU.
The diagram above shows why the unit is the KV head. With grouped-query attention, several query heads share one set of keys and values, so they also share one window assignment. They still compute different attention weights from different queries, and the model's trained projections learn how to combine what each head group saw.
Here is the arithmetic for one decode step (illustrative widths; the paper does not publish these exact values for its 128K benchmark). Take a 131,072-token context, 8 KV heads, a 64-token sink and a 4,096-token near window. The far past is 131,072 − 4,160 = 126,912 tokens, so each head's stretch is 126,912 ÷ 8 = 15,864 tokens. Each head reads 4,160 + 15,864 = 20,024 keys instead of 131,072, about 6.5× fewer. Summed over the eight heads that is 160,192 key reads instead of 1,048,576, and every one of the 131,072 tokens is still read by at least one head. A sliding window with the same 20,024-token budget per head would leave the oldest 111,048 tokens unread by every head.
| Pattern | What one head sees | Who covers the distant past | Cost to choose the pattern |
|---|---|---|---|
| Full attention | Every earlier token | Every head, duplicated | None |
| Sliding window | The last W tokens | No head | None |
| Routed sparse (MoBA, DSA) | Blocks or tokens chosen by scores | Whatever the router picks | A pooling or indexer pass over the keys |
| CoWA | Sink + near window + one far slice | Exactly one head per far slice | None: fixed by position and head index |
The coverage is what carries the quality, not the sparsity. In a window-matched ablation at 8K tokens, CoWA with full collective coverage scores 89.73% on associative recall (a test that asks the model to retrieve a value bound to a key stated earlier in the sequence) against 89.97% for full attention, while layouts that give several heads the same far window, at the same per-head width, score substantially worse. In training, the savings grow with context: at 14B parameters CoWA cuts total training FLOPs by 3.1% during 4K pre-training but by 28.5% during 32K long-context training, likely because attention is a larger share of the work at longer lengths.
One limit is easy to misread. The lower decoding memory the paper reports (about 7.6× lower peak operator memory, 8.4 MiB) is working memory for the attention computation, not the KV cache. The windows move forward as the sequence grows, so a token that is in one head's far slice now will fall into another head's slice later. Each GPU still stores the full KV history for its local heads, and the paper lists window-aware offloading of that cache as future work. For cache size, the relevant tools are still GQA and the methods in Memory Cost.
Goes deeper in: LLM Internals → Attention → Multi-Head Attention
Related explainers
- HydraHead — head-axis attention hybridization — another per-head design, which chooses full or linear attention head by head instead of splitting the history
- ConSA — controllable attention sparsity — learns where to place full and sliding-window attention, where CoWA fixes the pattern by position
- Reference Sliding Window Attention — a sliding window that does shrink the KV cache, at the cost of dropping old output tokens