LLM·

ThunderSyncRL — Gradient streaming without policy staleness — What does it mean?

The news. On October 5, 2026, researchers from Stanford University, Yonsei University, Korea University and Bespoke Labs posted ThunderSyncRL (arXiv 2610.05935). They trained Qwen-3.8-27B as a coding and terminal agent with GRPO and with on-policy distillation, on one node of eight NVIDIA B200 GPUs, and evaluated on SWE-bench Verified and Terminal Bench 4.0. The method reached the synchronous baseline's pass@3 targets up to 1.9× faster, and at a fixed 128-GPU-hour budget it scored 1.27 to 2.47 percentage points above asynchronous training. Read the paper →

Picture a teacher who grades a class's exam on a curve. A strict teacher reads no paper until the last student hands in, because no grade is final until the class average is known. When some students finish in ten minutes and others take three hours, that teacher sits idle for most of the exam. Synchronous agent training is the strict teacher: the learner GPUs compute no gradient until the slowest trajectory in the batch has finished and been scored. It is the same wait-for-the-longest problem as static batching at inference time, where one long sequence holds a whole batch. In the paper's GRPO run, each step spent 650.0 s on rollout and verification and then 646.4 s on learner work, one strictly after the other.

The obvious fix is to start grading next week's exam while this week's is still running, using last week's answer key. That is asynchronous RL: rollout and learning overlap across updates, but some trajectories were sampled by a policy that has since changed. This policy staleness can call for off-policy corrections or explicit staleness limits. The curve seems to block anything better. In GRPO, a trajectory's weight in the update is its advantage, its reward minus the group mean divided by the group's standard deviation, and neither number exists until the last trajectory in the group reports its reward.

ThunderSyncRL's key observation is that the curve is applied by multiplication, so the expensive marking can happen first and the curve last. Call one trajectory's gradient g, its reward r, and the group's mean and standard deviation μ and σ. GRPO treats the advantage as a constant while differentiating, so the group update Σ (r − μ)/σ · g rearranges exactly into (1/σ) · (Σ r·g − μ · Σ g). (This simplified form leaves out the paper's division by group size and the small ε it adds to σ for numerical stability; neither changes the argument.) The learner therefore runs each trajectory's backward pass as soon as its reward arrives and adds the result to two running tallies, Σ g and Σ r·g. When the group closes, μ and σ are two scalars and combining the tallies costs almost nothing.

After every group is combined, the system takes one optimizer step, and publishes the new weights to the rollout servers before the next batch begins. No trajectory is ever trained against a policy that changed while it was running. The paper's summary is that the method is asynchronous in execution and synchronous in policy updates.

The same rule, start each piece of work once its inputs are fixed, also covers on-policy distillation (OPD). A frozen teacher scores the student's tokens one agent turn at a time. The learner can run a turn's backward pass while that turn's tool call is still executing in the sandbox. That overlap fills the gap between model output and the next observation in one agent tick. The only quantity OPD cannot know early is the total number of scored tokens in the batch, which normalizes the loss, so ThunderSyncRL divides by it at the end. This is the same move continuous batching makes at inference: act on each sequence when it is ready instead of at a batch boundary. The paper proves both equivalences and confirms them numerically on B200 and H200 GPUs.

ScheduleWhen learner work startsPolicy stalenessPrice paid
SyncAfter the whole batch and every rewardNoneLearner idle 49.5% of each step in the paper's GRPO and OPD runs (paper, Fig. 3)
AsyncOverlaps the next rollout batchOne updateTrains on stale-policy data; may need correction or staleness limits
Fully AsyncNo batch barrier; weights change while trajectories runVariableTrains on stale-policy data; may need correction or staleness limits
ThunderSyncRLPer trajectory (GRPO) or per turn (OPD), once inputs are fixedNoneExtra memory for partial gradients: peak 117.0 → 122.8 GB per learner GPU at most (paper, §4.3)

Where the step time goes. Hold the paper's reference setup fixed: Qwen-3.8-27B, GRPO with 32 prompts × 8 trajectories = 256 trajectories per update, one eight-GPU B200 node. Rollout and verification take 650.0 s in both schedules. Synchronous training then adds 646.4 s of learner work, so its step is about 1,296 s before publishing weights. ThunderSyncRL has done almost all of that learner work while rollouts were still running: only 20.5 s remains after the last reward, and with publication its whole step is 687.7 s. 1,296 ÷ 687.7 ≈ 1.9× shorter steps at zero policy lag. Over a run this compounds: the GRPO target of 74.01% pass@3 on SWE-bench Verified arrived at 121.3 GPU-hours instead of 233.5.

The gain depends on balance: streaming can only hide learner work behind rollout work, so it pays most when the two take similar time. The 27B model is close to balanced. GLM-4.7-Flash-31B activates only about 3B parameters, so its learner work is short, rollout bounds the step, and there is less to hide. The paper also notes that objectives whose gradient inputs resolve early, such as turn-level OPD, leave more learner work to overlap than objectives that wait on group-level signals.

Goes deeper in: LLM Internals → Batching → Static Batching

Related explainers

Frequently Asked Questions

Check what you knowMap your AI & GPU knowledge across every track — free, role-based