Agent·

AutoCompact — Learned compaction timing — What does it mean?

The news. On October 1, 2026, Xuan Zhang and co-authors posted AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents (arXiv 2610.02163). It adds a compact() action to a terminal-based coding agent built on Qwen3-Coder-30B-A3B-Instruct, fine-tunes it on 1,052 judge-corrected trajectories, then runs reinforcement learning with task success as the only reward. It reports pass rates of 39.6% on SWE-bench Verified and 24.5% on SWE-PolyBench Verified, gains of 9.2 and 5.0 points over the same base model. Read the paper →

Wipe the board when a stage ends, not when it is full

Picture a team solving a problem on one whiteboard. Under the usual rule, nobody wipes anything until the board is full. By then the board is mostly crossed-out guesses, and the wipe happens wherever the marker happens to be — sometimes halfway through a calculation that still needed the numbers on the left. AutoCompact's rule is different: wipe the board when a stage is finished, and leave a short note in the corner that says what was found and what to do next.

For a coding agent, the board is the context window, the crossed-out guesses are search results, file dumps and failed attempts from exploration, and the corner note is the working state — a compact version of the agent's state object: conclusions so far, relevant code and workspace status, and remaining actions. The paper's own example is the clean case: once the agent has located the bug, the search history that found it is no longer useful, and the next stage (writing the fix) needs only the location and a plan.

Three decisions, not one

The paper splits compaction into three decisions. A length trigger answers the first one by counting tokens instead of reading task progress; the other two still have to be handled either way.

  • When to compact. It should follow task progress: compact after a stage is resolved, keep exploring while evidence is still needed.
  • What to keep. A summary that drops the target file name forces the agent to search for it again.
  • How to continue. The agent's next actions must follow the note, instead of re-running a search whose answer is already written down.

These line up with the context failure modes: exploration that stays in context after it stops being useful distracts the model, and a bad summary loses state the task still needs. The paper's evidence for the first point is that proactive compaction helped even when nothing was close to overflowing — every proactive method it tested beat the full-history agent in a 256K window.

Training it: corrections that are actually executed

The base model almost never calls compact() on its own, even when the prompt describes when to use it, so the authors had to show it examples. They ran the base agent on 379 training tasks and had a judge review each proposed step as it happened. The key detail is that the judge's correction runs in place of the original step, so the rest of the trajectory continues from the corrected state. Of the 1,052 trajectories collected, 24% teach a trigger correction (compact here, or keep exploring), 53% teach a working-state correction (fix what the summary records), and 23% teach a continuation correction (follow the note).

An earlier method, SWE-Compressor, inserted compaction calls into finished trajectories after the fact and kept the original later actions. Those examples never show the model what acting on its own summary looks like. With the same base model and a similar amount of data, AutoCompact's fine-tuned version scored 1.2 points higher on SWE-bench Verified and 1.6 points higher on SWE-PolyBench Verified.

Then reinforcement learning, with one reward

Supervised fine-tuning copies the judge's local choices; it does not check whether those choices led to a passing patch. So the second stage runs GRPO on SWE-Gym, a separate set of repository tasks with tests, with a binary reward: did the final patch pass the tests or not. Because each compact() call rewrites the context, a trajectory is cut into segments at every rewrite, and every segment shares the trajectory's single reward, so the decision to compact, the summary, and the moves after it are all credited or blamed by the final outcome. There is no separate reward for good summaries.

Even so, summaries got better. After fine-tuning the agent compacted on 44.3% of tasks; after RL, on 58.5%. Summaries missing key state fell from 3.1% to 0.2%, and summaries missing a next action fell from 8.2% to 2.2% (counted by keyword screening with manual spot checks).

MethodCompaction triggerTrainingContextSWE-bench VerifiedSWE-PolyBench Verified
Full history (base)None—256K30.4%19.5%
Fixed compactionLength—16K28.8%18.6%
CompactionRLLengthRL16K32.7%19.8%
SelfCompactRubric in the prompt—256K31.7%20.6%
SWE-CompressorLearnedSFT256K31.0%20.1%
AutoCompact-SFTLearnedSFT256K32.2%21.7%
AutoCompactLearnedSFT → RL256K39.6%24.5%

Every row uses the same base model and scaffold, averaged over three runs (paper Table 1). The 16K rows force compaction when the context passes 16K tokens; in the 256K rows no trajectory reached the forced threshold.

Worked example: where the 46 extra tasks come from

Hold two things fixed: the benchmark (500 SWE-bench Verified tasks) and the base model (Qwen3-Coder-30B-A3B in a 256K window). The reported pass rates are three-run averages, so multiplying by 500 gives task equivalents, not a list of specific tasks. The full-history agent solves 30.4%, or about 152 tasks. Fine-tuning on judge-corrected trajectories lifts that to 32.2%, or about 161 tasks — 9 more. RL adds another 7.4 points, about 37 more tasks, for 39.6%, or about 198 tasks. So of the roughly 46 extra tasks, about four in five (37 of 46) arrive in the RL stage — but RL trains coding and compaction together, so that share is not all credit for compaction. What the setup does rule out: no run in this window reached the forced-compaction threshold, so none of the gain comes from avoiding overflow.

Is it compaction, or just a better-trained model?

The authors ran the same trained checkpoint twice: once normally, and once with every compact() call skipped, so the history keeps growing. Actually executing compaction mattered most when money was tight: the normal run's pass rate was 19.9 points higher at a $0.10 per-task budget, and still 1.9 points higher at $4.00. Each run stops when its cost reaches the budget, so a plausible reading is that a run carrying its whole history pays to re-read stale tokens on every step and hits the limit sooner. (The cost figures price cached input tokens at 20% of the normal rate, the prompt-caching discount the authors assume.)

Two limits are worth stating. The results come from one 30B coding model and one agent scaffold (the loop and tools around the model), so other model families and harnesses are untested. And a summary can contain the right facts and still mislead: in one case study, the fine-tuned model's summary recorded an unresolved syntax error yet proposed finishing, and that run failed. The authors call the missing property summary self-consistency — the next action in the note must fit the state the note records. For a production harness, the paper's 16K test suggests keeping the length trigger as a fallback: with both active, the learned trigger still raised pass rates at every budget.

Goes deeper in: AI Agents → Context Engineering → The 4 Fixes

Related explainers

Frequently Asked Questions

Check what you knowMap your AI & GPU knowledge across every track — free, role-based