GPU·

ThunderEP on PCIe GPUs — Ring relay vs single-step host broadcast — What does it mean?

The news. On September 30, 2026, researchers at Seoul National University posted Efficient Expert-Parallel Communication on PCIe-Connected Consumer GPUs on arXiv. Their system, ThunderEP, replaces the NCCL All-Gather and Reduce-Scatter that vLLM uses for MoE dispatch and combine. On two machines with six RTX 4090 or six RTX 5090 cards each, they report average speedups of 2.00× for dispatch and 1.53× for combine over NCCL, and, on the RTX 5090 machine, up to 1.66× faster prefill (prompt processing) and 1.26× higher decode (token generation) throughput than vLLM, across Qwen3-30B-A3B, GPT-OSS-20B and GPT-OSS-120B. Read the paper →

Picture six offices on one floor with no doors between them. The only way to pass paper from one office to another is to walk it down to the shared mailroom, leave it in a pigeonhole, and have the other office walk down and collect it. That is a PCIe machine full of consumer GPUs: with no peer-to-peer path, every GPU-to-GPU byte goes into pinned CPU memory and back out, so it crosses PCIe twice. Data-center GPUs talk over NVLink or GPUDirect P2P instead, and the specialised MoE libraries the paper lists (DeepEP, pplx-kernels, NCCL EP and others) are built on that direct path, so the paper reports they cannot run here. vLLM, SGLang and Megatron-Core therefore fall back to NCCL.

NCCL moves an All-Gather with a ring: each GPU passes the shard it just received on to its neighbour, and after N−1 steps everyone has everything. In the office, that is a relay of runners. Office 0 drops its stack in the mailroom, office 1 picks it up and drops it again for office 2, and so on. With doors between offices, each hand-off is one short walk and the ring is close to the fastest schedule there is. Without doors, every hand-off is a full round trip through the mailroom, so the same stack is carried down and up again at every hop. The paper also measures that the mailroom gets crowded: on the RTX 5090 machine one GPU alone reaches 56.5 GB/s, but with a second GPU on the same PCIe host bridge each one drops to 35.5 GB/s.

ThunderEP's fix is to treat the mailroom as a notice board instead of a relay point. Each GPU posts its own shard into host memory exactly once, and every other GPU reads the N−1 shards it needs directly, so nothing is ever forwarded and the collective finishes in two phases however many GPUs take part. The ring moves 2(N−1) shard-loads of PCIe traffic per GPU; the single step moves N. The ratio 2(N−1)/N equals 1 at two GPUs, which is why the paper says the design only breaks even there and gains as the machine fills up.

Per GPU, All-Gather (dispatch)NCCL ringThunderEP single step
PCIe traffic, N GPUs2(N−1) shardsN shards
PCIe traffic, N = 6 (the paper's machines)10 shards6 shards
Sequential hops a shard can takeup to N−1 host round trips1 post + 1 read
Synchronization steps at N = 652
Sustained bandwidth at 1 GB, RTX 5090 (paper, §5.2)23.0 GB/s35.3 GB/s
Sustained bandwidth at 1 GB, RTX 4090 (paper, §5.2)20.8 GB/s26.5 GB/s

Cutting the walks is half the story; the other half is who does the walking. NCCL runs its copies as kernels on the SMs, the same cores the expert matrix multiplies need, so in the office the desk staff stop working to carry mail. ThunderEP hands the carrying to the GPU's DMA copy engines, the mail carts, so the SMs stay free to run the expert matrix multiplies (GEMMs) while tokens are in flight. The cart has limits: it only moves contiguous blocks and it cannot add numbers, so the final sum in combine still runs on an SM. Because a PCIe link is full-duplex, ThunderEP puts uploads on one stream and downloads on another and cuts each transfer into chunks, so one chunk goes down while the previous one comes up. The paper reports that NCCL, on machines without P2P, alternates the two directions chunk by chunk instead.

The last cost is the flags. A receiver must know a stack has fully arrived before it reads it, and on these machines the only place both GPUs can see is host memory, so every check of a completion flag is itself a trip across PCIe. NCCL's low-latency protocol stores a 4-byte flag beside every 4 bytes of data and polls with many threads. ThunderEP keeps the data contiguous, stores the flags separately, and polls with one thread per sender. It also lets a GPU start reading from each sender as soon as that sender's flag goes up, instead of waiting for the slowest one, and alternates between two host buffers, reusing each one only after the receivers report they have finished reading it, so a fast sender never overwrites data a slow reader is still collecting.

Worked example (illustrative traffic calculation). Hold two things fixed: N = 6 GPUs and an illustrative shard of 8 MB per GPU, in decimal units. The ring makes each GPU move 2(6−1) = 10 shard-loads, so 10 × 8 MB = 80 MB crosses that GPU's PCIe link for one dispatch. The single step moves 1 upload + 5 downloads = 6 shard-loads, so 6 × 8 MB = 48 MB: 40% fewer bytes for the same result. If you further assume the transfers run one after another at a steady 35.5 GB/s (the paper's figure for two GPUs sharing a host bridge), that is about 2.25 ms versus 1.35 ms. Treat those times as byte-count estimates, not measured latencies: ThunderEP also overlaps its upload and download directions, and real bandwidth depends on how many GPUs contend for the bridge. Combine is different: with the sum done on the GPUs, the paper's Reduce-Scatter moves the same 2(N−1) shard-loads as NCCL, which is why its large-message bandwidth is about equal (24.5 vs 23.2 GB/s on the RTX 5090) and its gain comes from shorter dependency chains on small messages.

Two engineering choices show where the design stops. During decode, vLLM replays each step as a CUDA graph, and a captured cudaMemcpyAsync keeps the address it was recorded with, so ThunderEP moves decode transfers back onto the SMs, where a kernel can choose between the two alternating buffers itself. The paper argues little is lost there, because decode transfers are only a few kilobytes and the SMs are not busy. For the same reason, overlapping communication with expert computation is used only in prefill; decode has too few tokens to form a pipeline. The authors also tried summing in host memory to save PCIe traffic and dropped it, because on their machines the CPU's DRAM bandwidth was lower than the combined PCIe bandwidth. They also leave routing-aware All-to-All, which would send each token only to the GPUs that need it, to future work: on these machines the CPU must read the router's decisions before it can issue each DMA, and that round trip cost more than it saved.

In the paper's breakdown for Qwen3-30B-A3B on the RTX 5090 machine (batch 24, 4,096-token prompts), ThunderEP cuts exposed MoE communication by 64.4% in prefill and 50.1% in decode, which lowers the whole Transformer layer by 38.9% and 15.2%. The ranks that take part are exactly the ones in vLLM's expert-parallel group; the router, expert placement and compute kernels stay unchanged. For a different MoE serving bottleneck, uneven expert load rather than slow wires, see SGLang's LPLB.

Goes deeper in: GPU & CUDA → Memory Hierarchy → NVLink & PCIe

Frequently Asked Questions

Check what you knowMap your AI & GPU knowledge across every track — free, role-based