LLM·

Sharpening Tax in Post-Training — pass@1 vs pass@K coverage — What does it mean?

The news. On October 1, 2026, researchers from Meta Superintelligence Labs, the University of Wisconsin–Madison and Stanford posted Sharpening Tax in Post-Training (arXiv 2610.01509). They compared 14 base/post-trained checkpoint pairs from four open model families (Gemma-4, Ministral-3, Qwen2.5 and Qwen3.5) on three tool-use benchmarks (BFCL v4 multi-turn, WebShop and an ACEBench subset of single-turn, multi-step and multi-turn tasks), 42 cases in total. Post-training usually raised one-shot accuracy while shrinking the set of tasks the model could solve at all given many tries. Read the paper →

Picture a locksmith facing a corridor of locked doors with a ring of keys. A base model, the network straight out of pre-training, is a huge jangling ring: on any one try the locksmith grabs a mostly random key, so the first try usually fails, but keep trying and most doors eventually open. RL post-training files the ring down to the few keys that opened the most doors during training. The first try now works far more often. The cost is the doors whose key was thrown away: within the paper's budget of 128 tries, many of them never open. The Sharpening Tax is the paper's name for that lost retry value: how much less extra attempts help once post-training concentrates the model's probability on a few rewarded behaviours.

Three numbers separate those two kinds of success. pass@1 is the chance one attempt succeeds, the number most leaderboards report. pass@K is the chance that at least one of K attempts succeeds, which measures solution coverage. passK is the chance that all K attempts succeed, which measures consistency; the pass@k vs pass^k step builds both formulas. The paper's central finding is that, for the larger models and large budgets it tested, post-training moves pass@1 and passK up while moving pass@K down: the model becomes more reliable on tasks it already handles and reaches fewer of the rest. Every attempt is sampled from the model's next-token distribution, so how spread out that distribution is, the property that temperature and top-p sampling control, decides how different the K attempts really are.

Why does the retry value disappear? The authors sorted every task by its outcome over 128 rollouts into three bins: always pass, pass given compute (solved on some tries, not others) and always fail. Post-training sharply shrinks the middle bin, the only bin where extra attempts help, and shifts those tasks toward the two extremes. In locksmith terms, the doors split into ones the favourite key opens every time and ones nothing left on the ring can open; the doors that open on the fifth or fiftieth try mostly vanish. These bins are observed outcomes within 128 tries, not proofs that a task is impossible. Model size moves the break-even point: for the smallest models sharpening often helps at small budgets, while on WebShop the sampling budget at which the Gemma-4 base model overtakes its post-trained version shrinks from more than 128 tries at 4B parameters to about 3 tries at 31B.

Task bin (WebShop, 128 rollouts)gemma-4-31B basegemma-4-31B post-trainedSource
Always pass0.0%26.0%paper §3
Pass given compute87.6%30.0%paper §3
Always fail12.4%44.0%paper §3

Here is how those bins become the headline coverage numbers. Hold the model (gemma-4-31B), the benchmark (WebShop) and the budget (128 rollouts per task) fixed. A task counts toward pass@128 if it is solved at least once, so pass@128 is always-pass plus pass-given-compute. For the base model that is 0.0% + 87.6% = 87.6%; for the post-trained model it is 26.0% + 30.0% = 56.0%, which matches the paper's reported "over 85%" versus 56%. Post-training bought 26 points of tasks the model now solves every time, and paid about 32 points of tasks it can no longer solve even once in 128 tries. The per-try arithmetic shows why a weak base model still covers so much (illustrative): a task the base model solves on 5% of tries has pass@1 = 5%, but pass@32 = 1 − 0.9532 ≈ 81%, while a task pushed to 0% stays at 0% for every K.

Measuring the tax, and paying less of it

To turn whole pass@K curves into one number, the paper defines a model's scalability as the gap between its pass@K ceiling and its pass@k curve, summed over every smaller k, then averaged over those K − 1 budgets and divided by the headroom 1 − pass@1. The Sharpening Tax is the base model's scalability minus the post-trained model's, so a positive tax means extra attempts help the post-trained model less, either because fewer tasks are ever solved or because success saturates within fewer tries. This calibrated tax was positive at K = 128 in 36 of the 42 model-benchmark cases, and a tax estimated from 8 rollouts predicted the ordering of the 32-rollout tax with a Spearman rank correlation of 0.85 (1.0 would be a perfect ordering), so it can be checked cheaply.

The obvious remedy, one higher temperature for every prompt at inference, does not fix it: in the paper's tests it usually lifted pass@K at large budgets but never improved pass@1, and sometimes lowered it. The paper's fix, posterior-tempered group sampling (PTGS), sets the temperature per prompt during RL training. It keeps a running estimate of each prompt's success rate as a probability distribution (a Beta distribution), draws one random guess from that distribution (Thompson sampling), and heats hard prompts toward τ to explore while cooling easy ones toward 1/τ to exploit, with τ between 1.2 and 1.5. The locksmith rummages deeper in the ring only at stubborn doors. Applied to PPO and GRPO training of Qwen2.5-7B-Instruct on the Sokoban and FrozenLake environments, the authors report that PTGS pays a smaller tax than fixed-temperature training while also improving pass@1.

If an agent system samples many attempts and keeps the one a checker accepts, pass@1 alone is the wrong number for choosing a checkpoint. In that setup the coverage curve decides which tasks are reachable, and these results suggest a large base model behind a light harness can beat its own post-trained version once the budget is high enough. If you serve one attempt per request, the post-trained model's higher pass@1 and consistency are what you want. Two limits on the evidence: the exact training recipes of the public post-trained checkpoints are not disclosed, and thinking mode was turned off for the main comparison.

Goes deeper in: AI Agents → Evals & Diagnostics → Pass/Fail vs Score

Related explainers

Frequently Asked Questions

Check what you knowMap your AI & GPU knowledge across every track — free, role-based