Filter overloaded MoE experts without changing token-side Top-K — Expert-to-token capacity filtering — What does it mean?
The news. On October 5, 2026, a team from Tongji University, Cornell, the Harbin Institute of Technology (Shenzhen) and the Shenzhen Loop Area Institute posted CIPHER-MoE: Balancing Efficiency and Routing Fidelity in Trillion-Scale MoE Training (arXiv 2610.05744). They fine-tune DeepSeek-V3, DeepSeek-V4-Flash, GLM-5 and DeepSeek-V4-Pro (284B to 1.6T parameters) on domain datasets and report 1.10×–1.94× faster training and up to a 64.9 percentage-point cut in the top-1 expert's workload, with domain-task scores close to plain MoE training. The paper says the code will be released soon. Read the paper →
Picture course registration. Every student lists the courses they want, and most of them want the same famous course. If that course has no seat cap, its one lecture hall overflows while the other halls sit half empty — and the term cannot start until the overflowing course has seated everyone. That is an MoE layer under expert parallelism. Each token lists its Top-K experts, a few popular experts collect most of the tokens, and every GPU waits for the device holding the hottest expert.
The skew is not small. On a domain-specific dataset, the paper measures that 10% of DeepSeek-V4-Flash's experts receive over 88% of the routing workload. A hot expert also needs memory for every token it holds, which is how the 1.6T-parameter DeepSeek-V4-Pro ran out of GPU memory before finishing a single step in the authors' setup.
The usual fixes act on the students' lists. An auxiliary load-balancing loss (GShard, Switch Transformer) or DeepSeek-V3's per-expert bias pushes the router toward even use, but it works on averages, so one micro-batch can still overload one expert. A fixed router that assigns experts by token position is perfectly even, but it stops matching tokens to experts by content; in the paper's Table 1 it trains faster yet scores lower on all four operations-research benchmarks than normal routing, which the authors attribute to weaker expert specialization when routing ignores content.
CIPHER-MoE leaves the lists alone and gives each course a seat cap. First the router runs exactly as before, and each token picks its Top-K experts. Then each expert ranks the tokens that picked it by cosine similarity between the token and the expert's router vector, drops any below a threshold, and keeps at most C of the rest — the best-fit students up to the seat cap. The paper sets C at 8 times the even share of tokens per expert. Most experts never reach that cap; the cap only bites on the hot experts, which is exactly where the step time and the memory peak come from.
What happens to a turned-away student matters as much as the cap. Rejection by one expert does not cancel the token's other assignments: each expert it picked filters independently, and the token still passes through the shared expert, the core lecture every student attends. It also stays in the training sequence and the loss. The paper compares two policies: strict drop, which sets that one assignment's routing weight to zero, and reroute, which sends it to another expert with spare seats and enough similarity. The paper recommends strict drop: it is faster in all 9 model-and-dataset pairs of its Table 2 and scores higher in 5 of them. The authors' explanation is that a rerouted token shifts what the receiving expert trains on, and its gradient can pull that expert away from its specialty — like seating a chemistry student in a literature seminar just because a seat is free.
Here is the arithmetic of one hot expert (illustrative numbers, not from the paper). Hold a micro-batch at 4,096 tokens, 256 routed experts, and K = 8. That is 4,096 × 8 = 32,768 assignments, so the even share is 32,768 ÷ 256 = 128 tokens per expert, and a capacity factor of 8 gives a cap of 1,024. Suppose the hottest expert is picked by 2,048 tokens, 16 times its even share, and at least 1,024 of them pass the similarity threshold. CIPHER-MoE keeps its 1,024 best-matching tokens and drops the other 1,024 assignments. That drops about 3% of all assignments (1,024 of 32,768) and halves the work on the busiest expert, from 2,048 token passes to 1,024. Halving that expert's work does not halve the step time, because communication, the other experts and the filtering itself also take time; the paper's measured end-to-end speedups are 1.10× to 1.94×.
| Approach | What it changes | Caps the busiest expert in each step? | Cost |
|---|---|---|---|
| Auxiliary balancing loss (GShard, Switch Transformer) | Adds a loss term that rewards even expert use | No — balances on average | Can conflict with the language-modelling loss |
| Loss-free bias balancing (DeepSeek-V3) | Adjusts a per-expert routing bias from past loads | No — reacts to history | No extra gradient, but a spike can still happen |
| Expert Choice routing (Zhou et al., 2022) | Experts pick tokens; tokens do not choose | Yes, fixed per-expert capacity | The token-side choice is gone |
| Fixed router | Assigns experts by token position | Yes, perfectly even | Can weaken expert specialization; NL4Opt 88.93 vs 93.08 for normal routing on DeepSeek-V4-Flash (Table 1) |
| Expert replication or migration (FasterMoE, FlexMoE) | Copies or moves hot experts to more devices | No — spreads the load across devices instead | Extra replicas, migration and runtime orchestration |
| CIPHER-MoE, strict drop | Tokens keep Top-K; each expert keeps its top-C tokens by cosine similarity | Yes, C = 8× the even share | Some assignments dropped; 1.13×–1.94× faster training (Table 2) |
The result is faster training at roughly matched domain-task scores, but not at zero cost on every benchmark. In its Table 2, strict drop on DeepSeek-V3 trains 1.94× faster on the operations-research data with a weighted average of 61.40 against 60.60 for plain training; on DeepSeek-V4-Flash it is 1.42× faster with 68.59 against 69.09. General-benchmark results are mixed: after operations-research fine-tuning, DeepSeek-V4-Flash's BBH score falls from 84.00 to 71.20 with strict drop, while most other benchmarks move by a few points either way (Table 4). All of the paper's experiments are domain-specific fine-tuning runs, not pretraining from scratch. On DeepSeek-V4-Flash, the filtering itself is reported to cost 1.57% of training time for strict drop and 3.85% for reroute.
Goes deeper in: LLM Internals → Transformer Block → The Feed-Forward Network
Where MoE fits among the other component swaps since 2017 is covered in LLM Internals → Transformer Block → Modern Variants & Scale.
Related explainers
- SGLang LPLB load balancing — the serving-side answer to hot experts: replicate them and rebalance load, instead of capping them
- SoftMoE differentiable routing — changes how the router picks experts, where CIPHER-MoE leaves the router's pick alone
- Manifold Power Iteration router alignment — improves load balance by reshaping the router's weights during pretraining