ITC-MoE — Adaptive Tucker compression of MoE experts — What does it mean?
The news. On 1 October 2026, a team from Hainan University, Xiamen University and Zhejiang Normal University posted ITC-MoE (Importance-guided Token-aware Compression). It shrinks the expert weights of MoE diffusion language models with an importance-weighted Tucker factorization, then adds two decode-time tricks for hot and cold tokens. On SDAR-30B-A3B-Chat-b32 it reports 96.33% accuracy on MultiArith (a math word-problem benchmark) at a 30% compression ratio, counted over the MoE layers only, and up to 7.22× end-to-end speedup; the implementation is built on the TEAM decoding framework, so not all of that speedup comes from compression alone. Read the paper →
Picture the bakery before the change: 128 recipe cards, each written out in full, every one listing its own chopping steps and its own glazing steps. Most of those steps are the same across cards. That is the situation inside an MoE layer: each expert is a feed-forward network with its own up, gate and down matrices, and the paper measures that the experts repeat a lot of each other's structure (Fig. 2c, a similarity score between expert outputs).
The usual way to shrink one matrix is to keep its strongest directions only, as SVD does. Doing that one expert at a time is like shortening each card on its own: every card still carries its own chopping and glazing steps. Tucker decomposition rewrites the whole recipe book at once. Stack the 128 expert matrices into one 3-D block, then find one shared input basis (the prep station), one shared output basis (the finishing station), a few core slices (the base batters) and, for each expert, a short list of mixing weights (the recipe card).
Two choices decide how good the rewritten book is, and they are where ITC-MoE differs from its Tucker baseline.
First, it measures which mistakes matter before it compresses. A calibration run records which input directions carry the most activation energy and which output directions move the loss the most, and the weights are re-scaled by both before factorizing. In bakery terms: keep the ingredients used in almost every order exact, and keep the flavours customers notice exact; the rest can be approximated. Removing this step drops MBPP from 59.92 to 45.91 at a 10% compression ratio.
Second, it learns how much to keep in each direction. The three modes are not equally safe to cut: the paper shows that truncating each one alone hurts the output by different amounts. Instead of a fixed split, ITC-MoE relaxes each rank into a smooth cut-off, trains the split to minimize output error under a fixed parameter budget, then hardens it. At a 20% ratio this lifts GSM8K (grade-school math problems) from 68.00 with fixed ranks to 79.38 with learned ranks.
The factorization also saves work at run time, not just storage. Because the prep station is shared, the bakery chops once per order instead of once per cake. Hold three things fixed: 8 active experts per token, the 3 matrices per expert (up, gate, down), and the 48 MoE layers of SDAR-30B-A3B. If each active expert ran its factored matrices separately, the input side of up and gate would run 8 + 8 = 16 times per layer and the output side of down 8 times, so 24 passes per layer. With shared factors, up and gate each run their input side once and down adds up the experts' contributions in the small core space before running its output side once: 3 passes per layer. Across 48 layers that is 1,152 → 144 shared-side passes per token. The per-expert parts (each expert's core transform, the up and gate output side, and the down input side) still run once per expert, so this is not an 8× speedup; the paper measures the shared-factor execution at up to 2.18×.
The second half of the method changes how decoding treats tokens: protect the ones about to be accepted, and narrow the expert choice for the rest. Hot tokens, the masked positions about to be accepted, get a small trained low-rank correction, because compression errors hurt them most; removing it drops MBPP from 59.92 to 51.75. Cold tokens route very alike, so they choose their 8 experts from a shortlist of 48 built from the cold tokens' average. In the paper's measurement cold tokens together touch about 17 experts versus 42 for hot tokens, and the 48-expert shortlist still contains the full top-8 set for 97.95% of cold tokens. That shortlist alone is reported at up to 13.0% speedup.
All of these results come from MoE diffusion language models (SDAR-30B-A3B and LLaDA2.0-mini). The Tucker idea itself does not depend on diffusion, but the paper does not test it on ordinary autoregressive MoE models, and it notes that a gap to the full model remains.
| Method | What it does | MBPP | MultiArith | Source |
|---|---|---|---|---|
| Original (uncompressed) | Reference model | 65.76 | 98.67 | Table 2 |
| D2-MoE | Low-rank method built for autoregressive MoE | 34.24 (30% ratio) | 98.00 (30% ratio) | Table 2 |
| TD-MoE | Tucker across experts, ranks set in advance | 12.45 (30% ratio) | 73.83 (30% ratio) | Table 2 |
| ITC-MoE | Importance-weighted Tucker, learned ranks, hot/cold handling | 47.47 (30% ratio) | 96.33 (30% ratio) | Table 2 |
Where it fits in the bigger picture: quantization shrinks a model by storing each number in fewer bits; Tucker compression shrinks it by storing fewer numbers, using structure the experts share. The two are separate levers, and some of the modern compression methods combine them. Compressing the expert bank as one object lets the experts store what they have in common once, which compressing them one by one cannot do.
Goes deeper in: LLM Internals → Transformer Block → The Feed-Forward Network
Related explainers
- SigmaScale: learned scaling for SVD compression — the one-matrix-at-a-time version of low-rank compression
- MAESTRO: pruning MoE experts by routing traffic — shrinks an MoE by removing whole experts instead of factorizing them