Retrieve long-range blocks with prefix-state hybrid attention — Prefix-state block retrieval vs sliding-window hybrid attention — What does it mean?
The news. On October 6, 2026, researchers from the Hong Kong University of Science and Technology (Guangzhou), with an independent researcher, posted PHBA: Prefix-State Hybrid Block Attention to arXiv. They train 760M- and 1.3B-parameter models from scratch on FineWeb-Edu at an 8K context, compare them with a Transformer, MoBA, GSA, KDA (a gated linear-attention model) and NHA trained on the same data, and report the strongest RULER and NIAH averages among the compared 1.3B models at 16K, 32K and 64K tokens, beyond the training length, plus the strongest LongBench average on real-document tasks with inputs capped at 8K. The paper also describes a Triton implementation; the code is promised on acceptance. Read the paper →
Picture catching up on a long series before the finale. A sliding-window hybrid watches only the last few episodes in full and relies on one season summary for everything older — fine until the finale turns on a detail from episode 3 that the summary blurred. Pure block retrieval, as in MoBA, lets you jump straight to episode 3 by scanning the episode guide, but you watch it cold, with no idea what led up to it. PHBA does both: you jump to the few episodes that match your question, and each one opens with its own Previously on recap of everything before it.
In attention terms, the series is the context, cut into blocks of 128 tokens. To route, PHBA represents each earlier block by the mean of its keys (its synopsis in the episode guide), scores the query against those means, and keeps the top-scoring blocks. The current block is always kept under ordinary causal attention, so with a budget of K blocks, K − 1 go to routing. Once a block is picked, the query reads its original token-level keys and values; the mean key was only used to choose it.
The recap comes from a gated slot memory: 128 key/value slots that a learned gate updates token by token. PHBA saves this memory at every block boundary, just before the block starts, so each saved state summarizes everything earlier and nothing from the block itself. When a query picks a block, it also attends to that block's saved state, so the evidence never arrives without its context. The current block's state covers all history before the local block. In the paper's ablation, the learned gate beats both an ungated recurrence and a plain mean of earlier tokens, most clearly beyond the training length.
The exact tokens and the recap slots compete inside one shared softmax, so each query decides how much weight goes to precise evidence and how much to compressed context. The two kinds of memory are scored in separate passes (the token side adds position information to its scores; the slot side carries none) and combined with the online-softmax merge from FlashAttention: each pass keeps a running maximum and sum, and merging them gives exactly the result of one softmax over both, without ever building the concatenated score matrix. The Triton kernels also read the selected blocks and states in place, from their original memory layout, instead of first copying them into a packed buffer.
Here is what one query sees in the paper's 1.3B setting. Hold the context at 8,192 tokens, the block size at 128, and the exact budget at 1/8 of the context. That gives 8,192 ÷ 128 = 64 blocks and a budget of K = 8 blocks: the current block plus 7 routed ones, or 8 × 128 = 1,024 exact tokens. Each of those 8 blocks brings a 128-slot prefix state, so up to another 1,024 slot entries join the softmax. A query at the last position therefore scores, per head, at most 2,048 token-and-slot entries instead of the 8,192 a dense layer scores — about a quarter — and can still reach any earlier block (our arithmetic from the defaults; earlier positions see fewer tokens, and the first block has no prefix state). The paper does not report inference memory, but the same defaults point at a cost worth knowing: keeping a state at all 64 boundaries, so that any block can be routed, means 64 × 128 = 8,192 saved key/value slot pairs per head per layer, as many entries as the token KV cache itself (our arithmetic from the defaults, not a measured figure).
| Design | Exact attention covers | Older history | RULER avg, 8K (1.3B) | NIAH avg, 8K (1.3B) |
|---|---|---|---|---|
| Full attention (Transformer) | every earlier token | nothing compressed | 15.9 (Table 1) | 44.5 (Table 1) |
| Gated linear memory (GSA) | none | one fixed-size gated state | 10.4 (Table 1) | 16.5 (Table 1) |
| Block retrieval (MoBA) | current block + top routed blocks | selected blocks stay exact; unselected blocks are invisible | 10.8 (Table 1) | 23.6 (Table 1) |
| Sliding-window hybrid (NHA) | recent tokens in a fixed window | one compressed recurrent state | 15.1 (Table 1) | 27.4 (Table 1) |
| PHBA | current block + top routed blocks | a prefix state for each block it attends to | 23.9 (Table 1) | 29.5 (Table 1) |
The gains come with a modest compute bill, and dense attention is not beaten everywhere. At 1.3B parameters, PHBA has the best language-modeling score of the compared models (WikiText perplexity, lower is better: 15.21 against the Transformer's 15.67) and the best RULER average at 8K, but the dense Transformer still leads the needle-in-a-haystack average inside the training length (44.5 against 29.5). On one NVIDIA H20 GPU, adding the prefix states incurs 6.7–11.9% overhead relative to MoBA (block size 256, 2K–16K tokens). Under the matched 1/8 budget, PHBA trains 18.4–34.6% faster than NHA at 2K, and that NHA configuration ran out of memory from 4K onward in the authors' implementation. With K = 4, PHBA reaches 8.79K tokens/s at 16K, 12.8% above full attention.
Treat this as a design result at small scale, not a production recipe yet. The models stop at 1.3B parameters and 100B training tokens with an 8K training context, the long-context tests stop at 64K, and the code is not yet public. The lasting idea is the question PHBA answers: when a layer keeps only a few exact pieces of a long history, what does it do with the gap in front of each piece — drop it (block-sparse), keep only the recent end (sliding window), or attach a summary of it (prefix states).
Goes deeper in: LLM Internals → Self-Attention → Computing Attention Scores
Why those skipped tokens matter for memory is covered in LLM Internals → KV Cache → Memory Cost.
Related explainers
- MiniMax Sparse Attention (MSA) — block-sparse attention in a shipped 1M-context model, where unselected blocks are simply skipped
- Head-axis attention hybridization — a different hybrid split: full versus linear attention per head instead of per layer
- Attention dilution — why exact long-range retrieval fails as the context grows
- Training-free linear attention — the compressed-state side of the tradeoff on its own