Parallelism trade-offs paper — Tensor vs pipeline parallelism for prefill and decode — What does it mean?
The news. On October 4, 2026, two researchers from Dell's Office of the CTO posted Characterizing Parallelism Strategies in LLM Inference: Fundamental Compute-Communication Trade-offs. They write the latency of TP, PP and hybrid layouts as equations, then profile Qwen2.5-32B-Instruct and Qwen2.5-14B-Instruct in vLLM v0.15 on 8× A100 80 GB GPUs joined by NVLink, with 1,000-token prompts, 200-token outputs and batch sizes from 1 to 128. Read the paper →
Picture the two kitchens. In the huddle kitchen, four cooks stand around every dish. Each cook chops a quarter of the ingredients, and then the team huddles to put the quarters together before the next step. Each cook does only a quarter of the work, but the team must huddle after every step, and when the order is small the huddle is mostly fixed start-up and waiting time, not food being passed around. That is tensor parallelism. The huddle is an AllReduce, and in the standard TP layout it happens twice in every transformer layer for every generated token. The paper measures what the huddle costs: at higher TP degrees, communication takes about 30–40% of decode latency, and in decode the AllReduce reaches only about 150 GB/s, roughly a quarter of the 600 GB/s NVLink rating. The messages are small, so fixed start-up cost dominates rather than raw bandwidth.
In the assembly-line kitchen, each station owns one block of layers and passes the plate to the next station. Stations never huddle; they only hand a plate to their neighbour, which is a cheap point-to-point transfer. The problem is the empty line. When only one plate is moving, three of the four stations stand idle, and that idle time is the pipeline bubble. The paper writes the idle share as (P − 1) / (W + P − 1), where P is the number of stages and W is the units of work entering the pipeline. The fill-and-drain cost is a fixed P − 1 stage intervals, so it only disappears when W is large.
The two phases of inference give the line very different amounts of work, and that is the whole result. Prefill pushes the entire prompt through at once, so in the paper's model W is the batch size times the prompt length, and the line stays full. Decode produces one token per request per step, so W is just the batch size, and at small batches the line is mostly empty. This is the same split the roofline model draws between compute-bound prefill and memory-bound decode, seen from the network side. The paper's measurements follow the model: PP-heavy layouts give the lowest time-to-first-token, and TP-heavy layouts give the lowest inter-token latency. With all 8 GPUs, the order for decode ran from PP8 (worst) through PP4TP2 and PP2TP4 to TP8 (best); for prefill at batch 128, PP-heavy layouts finished first, because TP's per-layer AllReduce grows with prompt length while PP only ships activations once per stage boundary.
| Layout (8 GPUs) | What each GPU holds | What crosses the wire | Main idle cost | Wins in the paper's runs | Source |
|---|---|---|---|---|---|
| TP8 | 1/8 of every layer | 2 AllReduces per layer, per token | ~30–40% of decode latency is communication (at higher TP) | Decode (ITL) | §3.2 |
| PP8 | 1/8 of the layers, whole | 1 activation hand-off per stage boundary | Pipeline bubble, large at small decode batches | Prefill (TTFT) | §3.3 |
| PP2TP4 / PP4TP2 | TP inside a stage, PP across stages | Both kinds of traffic | In between; the bubble share depends only on PP depth and work, not on TP degree | Intermediate on both | §2.3, §3.4 |
Here is the bubble formula with the paper's own setup plugged in (our arithmetic on the paper's equation, not a measured figure). Hold the layout at PP8, so P = 8 and the fixed fill-and-drain cost is 7 stage intervals. In decode with batch 1, W = 1, so the idle share is 7 / (1 + 7) = 87.5% — in the paper's pipeline model, seven-eighths of the stage time is idle. Raise the decode batch to 64 and W = 64, so the idle share falls to 7 / 71 ≈ 9.9%. Now take prefill of one 1,000-token prompt: the paper's model counts W = 1 × 1,000 = 1,000, so the idle share is 7 / 1,007 ≈ 0.7%. The same eight-stage line wastes almost nothing on a banquet order and almost everything on a single plate, which is why decode at small batches is where TP pays for its huddles.
One practical reading follows from this, although the paper does not test it. If prefill wants an assembly line and decode wants a huddle team, a system that runs the two phases on separate GPU pools can choose a different layout for each. That is one argument for prefill/decode disaggregation. Inside a single engine, the choice is made by how ranks are laid out across the tensor, pipeline and data-parallel axes and which process groups exchange data. The model's limits are worth stating too: it was validated on one 8-GPU NVLink node with dense Qwen2.5 models, so cross-node links, mixture-of-experts layers and very different prompt lengths can move the crossover points. Measure the result as TTFT and ITL against your SLOs, not as raw throughput.
Goes deeper in: LLM Serving → Prefill/Decode Disaggregation → Full Disaggregation
Related explainers
- Batch-sharded sampling across tensor-parallel ranks — another place where TP's per-step coordination cost shows up
- Revocable decode-node leases — how a disaggregated cluster shares decode capacity once the phases are split