TopK-Guided — Budget-aware activation sparsity — What does it mean?
The news. On October 1, 2026, Mukund Agarwalla and Chih-Jen Lin of MBZUAI posted TopK-Guided: Adaptive, Budget-Aware Activation Sparsity for Efficient LLM Inference, a workshop paper for AXIOM 2026. It is a training-free change to an existing method, WINA, in exactly two places: how each token's mask is chosen, and how the sparsity budget is split across transformer blocks. On Llama-2-7B and Llama-3-8B it reports the lowest perplexity and highest average accuracy in all twelve comparisons against TEAL and WINA, at essentially the same sparsity-dependent projection compute as WINA. Read the paper →
Picture the mixing desk again, one desk per weight matrix. Every token that passes through a transformer block hits a handful of big projections — attention's Q, K, V and O, and the feed-forward network's gate, up and down matrices — and each entry of the input vector is one channel on the desk. Activation sparsity is the decision, made fresh for every token, to mute the quietest channels: zero the entries that matter least for this token, and the weight column each one multiplies never has to be touched. Zero a fraction s of the entries and you skip a fraction s of that projection's multiply work. The model itself does not change. That is the difference from weight pruning, which unplugs the same amplifiers for every song; the LILA explainer covers that static version.
Why bother? During decode, one token's vector meets the whole weight matrix, so the GPU spends most of its time pulling weights out of memory rather than doing math. A muted channel is a column a kernel could, in principle, leave unread. That makes the muted fraction the knob that sets the saving — and a method that cannot hit the fraction it promises cannot promise a speedup either. The paper sparsifies during the decode phase only, and it counts its savings in FLOPs rather than measured time; a kernel that actually skips the reads is listed as future work.
The two existing rules sit at opposite ends of the desk. TEAL draws a fixed volume line per projection, calibrated offline so that on average a fraction s of activations fall below it, and mutes every channel under the line. A quiet song loses many channels and a loud one few — the paper calls this token adaptivity — but because the line is frozen while songs keep changing, the realised share drifts away from the target. WINA instead always mutes the same count: it keeps the k entries with the largest weight-informed score (the activation's size times its column's length) and zeroes the rest. WINA always hits the target exactly, but a token that is mostly silence gets the same number of mutes as a token where every channel matters.
TopK-Guided's first change keeps TEAL's volume line only as a guess, then forces the count into a narrow band. For each decode token it counts the fraction of entries below TEAL's calibrated threshold — the token's natural sparsity — clips that fraction to within ±δ of the target, and runs WINA's TopK with the resulting count. With δ fixed at 0.01, every token's realised sparsity stays within one percentage point of the target its projection was assigned after the block and per-matrix split (up to rounding to a whole entry), and WINA is simply the case δ = 0. The extra work per token is one comparison pass and a cutoff search done with quickselect instead of a partial sort, which the paper calls negligible next to the projection itself.
The second change is about the chain of desks. TEAL and WINA hand every transformer block the same budget s, yet the paper measures that early blocks are far more sensitive than later ones. So TopK-Guided runs one calibration pass, sparsifies each block on its own at a probe level, records how far that block's output moves, and then shifts mutes from sensitive blocks to robust ones. Before clamping, the budgets still average to s; the clamp can nudge that average. The rule is si = s + λ(ē − ei), clamped between 0.05 and 0.95, where ei is block i's measured output change, ē the mean over all blocks and λ how hard to redistribute. Each block's budget then feeds WINA's existing greedy split across that block's seven weight matrices, unchanged.
A worked example — one projection with d = 4,096 input entries (illustrative; the paper does not fix d, and this ignores the per-block and per-matrix split). At a target of s = 0.5, WINA keeps exactly ⌊4,096 × 0.5⌋ = 2,048 columns for every token. TopK-Guided with δ = 0.01 may keep anywhere from ⌊4,096 × 0.49⌋ = 2,007 to ⌊4,096 × 0.51⌋ = 2,088 — a swing of −41 to +40 columns, roughly 2% of the kept work, decided per token. TEAL's line is the one that slips: on Llama-2-7B at the same 50% target, the paper measures TEAL's realised sparsity at 48.36% on C4 and 47.64% on cc100. On C4 that means keeping 51.64% of 4,096 ≈ 2,115 columns on average — about 67 more than promised, roughly 3% more projection work than the 2,048 budget — and the overshoot changes with the dataset. For the block budgets, take illustrative sensitivities: λ = 0.4, a mean ē = 0.2, a fragile early block at e = 0.5 and a robust late block at e = 0.1. The early block runs at 0.5 + 0.4 × (0.2 − 0.5) = 0.38 and the late one at 0.5 + 0.4 × (0.2 − 0.1) = 0.54.
| Method | Mask rule | Count per token | Block budget | C4 perplexity at s = 0.7, Llama-2-7B / Llama-3-8B |
|---|---|---|---|---|
| TEAL | Fixed threshold, calibrated offline | Varies; realised 48.36% at a 50% target on C4 | Same s everywhere | 15.85 / 24.09 |
| WINA | TopK on activation size × column length | Exactly k, every token | Same s everywhere | 10.72 / 27.78 |
| TopK-Guided | Threshold guess, clipped to ±0.01, then TopK | Moves within ±1 point of its assigned target | Shifted by measured block sensitivity | 10.15 / 21.55 |
The gain is tiny where little is muted and grows where most is. At s = 0.3 the three methods sit within 0.03 perplexity of each other on Llama-2-7B; at s = 0.7 the table above shows TopK-Guided ahead on both models, with the widest gap on Llama-3-8B. That is the desk logic: when most channels stay on, which ones you mute barely matters, but when seven in ten go silent every choice of channel and of desk counts. The paper's ablation adds each change to WINA on its own: the masking change alone meets or beats WINA everywhere, while the block budgets alone pay off mainly at s = 0.7 and slightly worsen C4 perplexity at s = 0.5 (7.35 against 7.34); together they are best at every level. The authors also name their limits plainly — models only up to 8B parameters, and savings counted as FLOPs rather than measured as wall-clock time.
Goes deeper in: LLM Internals → Transformer Block → The Feed-Forward Network
Related explainers
- LILA: calibration-free neuron pruning — the static cousin: remove the same feed-forward neurons for every token instead of choosing per token.
- FlexEE: KV-compatible early exit — another way to skip decode work per token, by stopping partway up the layer stack.