Agent·

Microsoft Agent Lightning 1.0 — Proxy-captured harness rollouts — What does it mean?

The news. On October 7, 2026, Microsoft Research released Agent Lightning 1.0, an open-source framework of about 3,500 lines for training existing agents with reinforcement learning while their real harness stays in control. In the post's coding example, about 6,000 training samples from the SWE-smith dataset, run through mini-SWE-agent, raised Qwen3.5-9B from 41.8% to 56.4% Pass@1 on SWE-bench Verified. Agents run as local processes or standard Kubernetes jobs and reach the model through an OpenAI-compatible proxy. Read the release → · Code on GitHub

Picture a driving school that never lets you drive your own car. It builds a simulator that looks like your car, but the mirrors sit in a different place and the brakes bite at a different point. You pass the course, and then you drive a car you never practised in. Agent RL usually works the same way. Most frameworks assume the trainer owns the interaction loop, so that each rollout is one continuous stream of tokens, and so developers rebuild the agent inside the trainer. But a real harness brings its own context management, tool protocols and execution logic. That is everything a harness adds around the model, and every part of it shapes what the model learns. In the post's words, the rebuild is costly, and the trained agent may not be the one that is deployed.

Agent Lightning keeps your own car and puts an instructor in the passenger seat. The only change to the harness is its model endpoint: it points at an OpenAI-compatible proxy, which forwards each call to the model being trained and records the prompt, the response and the log probabilities, linked to the rollout they belong to. The harness keeps doing its normal work: managing context, calling tools, starting subagents and recovering from errors. The proxy's log resembles the one-span-per-model-call trace that production observability records, plus the log probabilities that training needs.

The instructor's notebook has a problem: it holds text, and RL needs tokens. A harness stores its context as text, but the RL update needs the exact token IDs the model sampled, and running that text back through the chat template and tokenizer can move the token boundaries. For example (illustrative), a model may have sampled read then _file, while the tokenizer, given the text read_file, produces read_ then file. Byte-pair encoding can split the same text in more than one way, so the split a tokenizer picks need not be the split the model sampled. When the boundaries no longer line up, adjacent model calls cannot always be merged into one sample. Subagents and context summarization split a run in the same way. So one drive can fill one notebook page or five, and the page count reflects how the harness behaved, not how well the model drove.

ProblemWhat goes wrong with a real harnessAgent Lightning's answerSource
RetokenizationText re-encoded into tokens may not match the sampled token IDs, so calls cannot always mergeHandled when samples are assembled (the post names the problem but does not detail the method)MSR blog
Variable sample countsOne rollout becomes 1 to N samples (retokenization, subagents, summarization)Keep every sample linked to its rolloutMSR blog
AdvantageA per-sample baseline counts long rollouts several timesRollout-level advantageMSR blog
Loss normalizationPer-sample averaging gives more weight to runs that produced more samplesRollout-level normalizationMSR blog
SchedulingSample count and length are known only after the harness finishes, but GPU count and parallelism are fixedA trainer built on verl (an open-source RL training library) maps the variable workload onto fixed resourcesMSR blog

Here is why the grade must be per drive (illustrative numbers). Take two rollouts on the same task. Rollout A succeeds (reward 1), and because its harness summarized the context partway through, it produces 4 samples. Rollout B fails (reward 0) and produces 1 sample. Scored per rollout, the baseline is (1 + 0) / 2 = 0.5, so A gets an advantage of +0.5 and B gets −0.5. Scored per sample, the baseline is (1 + 1 + 1 + 1 + 0) / 5 = 0.8, so each of A's samples gets only +0.2 and B's sample gets −0.8. The success now looks barely better than average because the harness happened to split it into more pieces. Normalization has the same problem. Averaged per sample, A carries 4 of 5 shares of the loss (80%), while averaged per rollout each run carries 50%. Per-sample statistics let the harness's bookkeeping decide how much each attempt counts; rollout-level statistics put the weight back on the outcome.

The instructor also has to change the car's settings without stopping the drive. In collocated async RL, rollouts and model updates share the same GPUs: when enough rollouts are collected, the gateway pauses new requests, lets in-flight calls finish, updates the model and resumes, and the harness never sees the switch. The post reports this as about 2× faster end to end than synchronous RL, using fewer GPUs than conventional asynchronous RL (its timeline figure draws four GPUs against eight). It does not give the GPU models, hyperparameters or total compute behind that figure. A rollout controller starts each agent run as a local process or a standard Kubernetes job and reports status back to the gateway, which the post presents as a way to avoid paid sandbox services.

On results, the post attributes the full 41.8% → 56.4% Pass@1 gain (+14.6 points on SWE-bench Verified) to RL training alone, with about 6,000 samples, and its ablation over 200 training steps shows rollout-level advantage plus rollout-level normalization reaching a higher validation reward than either rollout-level advantage alone or sample-level advantage. The post does not report other harnesses, other base models or other benchmarks, so treat the gain as one configuration, not a general rate. A single pass@1 number also does not show run-to-run variance.

Goes deeper in: AI Agents → The Agent Loop & State → The Anatomy of a Harness

Related explainers

Frequently Asked Questions

Check what you knowMap your AI & GPU knowledge across every track — free, role-based