Denoising Surface paper — Denoising Workload Surface — What does it mean?
The news. On September 30, 2026, a team led by researchers at Wuhan University posted Denoising Surface: Modeling and Predicting Inference Cost for Diffusion LLM Serving. Serving LLaDA2.0 models in SGLang 0.5.10 on NVIDIA PRO 6000 GPUs, their predictor cut cost-prediction error by up to 2.50× compared with scalar predictors. A shortest-job-first scheduler driven by it cut average end-to-end latency by up to 1.92× and time-to-first-token by up to 2.22× compared with first-come-first-served on LMSYS-Chat-1M chat traffic. Read the paper →
Picture the bakery manager with one oven and a line of orders. Baking the smallest order first keeps the line moving, but only if the manager can tell which order is smallest before baking it. At a normal bakery that is easy: count the cookies, because every cookie takes one trip through the oven. That is how autoregressive serving sizes a request. One forward pass produces one token, so a predicted output length is a predicted cost, and shortest-job-first (SJF) schedulers built on that prediction cut the time short requests spend waiting in the queue behind long ones.
A diffusion LLM breaks the bakery rule. It writes its answer in blocks of 32 tokens (LLaDA2.0's default) and finishes each block with repeated denoising steps, which are forward passes that each reveal one or more of the block's masked tokens. An easy block can finish in a pass or two; a hard one can need close to 32. So the cookie count stops tracking oven time: the paper measured requests with the same input and output lengths whose inference times differed by up to 11.47×.
Counting oven passes (total denoising steps) is closer, but it misses a second effect: passes inside a block are not equal in cost. On LLaDA2.0-mini, a mixture-of-experts model, the last denoising step of a block averaged 8.62 ms against 6.04 ms for the first. The attention time barely changed. The extra time came from the expert layers: as the block's hidden states evolve, its tokens spread over more experts (32.7 → 66.3 activated experts per layer on average), and the MoE kernel pads each expert's rows separately.
The fix is the manager's shaded grid. DWS keeps the work as a two-dimensional surface, one row per output block and one column per denoising step inside that block, and shades each cell with the probability that this request will actually run it. A cell (block b, step s) is active only if generation reaches block b and that block needs at least s steps. So each cell is the product of two chances: that the answer is still running at block b, and that block b, once reached, still needs work at step s. The first can only fall as b grows and the second can only fall as s grows, so every row fades to the right and the first column fades downward. A later block can still need more steps than an earlier one.
Then the stopwatch. The predicted cost is the surface multiplied cell by cell with a cost-per-cell matrix profiled on the actual deployment, plus linear terms for prefill (per prompt block) and KV refresh (per finished output block). On 20K real LMSYS-Chat-1M traces, the cell-cost model with a prompt-length correction fit measured denoising latency with R² = 0.9956 (1.0 would be a perfect fit).
The surface is predicted before the request runs, from the prompt alone. A small sentence encoder (all-MiniLM-L12-v2) feeds two heads: one for how many blocks the answer will use and one for how many steps each block will take. Training goes coarse-to-fine: first on rough targets learned together (the block count, and step counts capped at several depths), then on the full surface. Quantized to INT8, the predictor runs on one CPU core, so it never competes with the model for the GPU. Because the surface describes the request and the stopwatch describes the hardware, a new GPU needs a fresh profile, not a retrained predictor, as long as the model and its decoding settings stay the same.
| Cost proxy | What it counts | What it misses | Measured on LLaDA2.0-mini |
|---|---|---|---|
| Output length | Tokens generated | One pass can reveal several tokens | Same input and output length, up to 11.47× spread in time (paper §1) |
| Total denoising steps | Forward passes | Late passes in a block cost more | Denoising-only DWS model: up to 1.80× lower error than step count (paper §6.2) |
| DWS | Probability of each block-step cell × its profiled cost | Still an estimate made before the request runs | Full predictor: up to 2.50× lower cost error than scalar predictors (paper abstract) |
Here is why the ordering matters (illustrative: denoising cost only, prefill and KV refresh left out). Take two requests that each run 64 denoising steps. Assume full 32-token blocks. Request A writes 32 blocks that each finish in 2 passes, which is 1,024 tokens. Request B writes 2 blocks that each need all 32 passes, which is 64 tokens. Output length says A is 16× bigger; step count says they are equal. Now price the cells with the paper's two measured endpoints. A only runs first and second passes; if the second costs about the same as the first, that is 64 × 6.04 ≈ 387 ms. If B's cost per pass climbs evenly from 6.04 to 8.62 ms across each block, its passes average about 7.33 ms, so 64 × 7.33 ≈ 469 ms. The request that looks 16× larger by token count is the cheaper one to denoise, and B costs about 21% more, so a length-based SJF queue would serve the wrong one first.
SJF has limits. Under light load every scheduler in the paper performed about the same, because requests rarely wait; the gap opened as the request rate rose. Always picking the cheapest job can also starve an expensive one, so the paper's scheduler discounts a waiting request's predicted cost by 10% every 30 seconds until it is picked. A cost predictor only pays off when there is a queue to reorder, which is exactly when time to first token hurts most.
Goes deeper in: LLM Serving → Inference Engine → The Scheduler