UserProxyBench audits simulated users — Simulated-user fidelity — What does it mean?
The news. On September 29, 2026, Ashish Jain (Sarvam AI) and Armaan Sandhu (UMass Amherst) posted UserProxyBench, a workshop paper that adds an evaluation layer on top of τ²-bench. They froze GPT-5.5 as the agent, swapped in seven different user proxies, and ran all 375 enterprise tasks across four domains (Telecom 114, Airline 50, Retail 114, Banking 97). Only the simulated user changed, yet the agent's mean task reward moved from 0.644 to 0.796, and 24.4% of the episodes the agent "won" contained a violation of the user's own script. Read the paper →
Picture the exam room again. The school wants to know whether the student can take a patient history: ask the right questions in the right order and reach the diagnosis. So the actor gets a script card that says, in effect, answer what you are asked and nothing more. Now one actor walks in and opens with every symptom, the test results and the likely cause. The student writes down the correct diagnosis and the examiner marks a pass. The pass is real, but it no longer measures history-taking, because the actor did that part of the job.
Interactive agent benchmarks have exactly this structure. In τ-bench a support agent must look up a customer, follow a policy and call tools, while a second LLM plays the customer from a private blueprint. The standing simulator guidelines shown to every proxy include an explicit rule: disclose information progressively and wait for the agent to ask. One Telecom proxy in the paper opens with "I'm John Smith, by the way. My number is 555-123-2002" before the agent has asked for either. The task still ends in the right state, so the agent still scores a pass. The benchmark only ever graded the student; nobody graded the actor. That is a case of the first item in the four eval failure modes: a part of the measuring instrument that was never itself validated.
UserProxyBench adds the second examiner. For each task, a model (Claude Opus 4.8) writes 4 to 8 atomic criteria from the blueprint, each one quoting the exact line of the script it enforces and each one about the user only, for example "does not give the phone number until the agent asks for it". A different model (GPT-5.6-Sol) reviews every rubric for grounding and scope. Across the 375 tasks that gives 2,141 criteria in four families: groundedness (the user does not invent device state or claim actions it did not take), premature disclosure, goal deviation (the user gives up or accepts an outcome the script forbids) and missed information (the user withholds something it was asked for). An episode's fidelity is 1 only if every applicable criterion passes, and the User Fidelity Score is the share of episodes that reach 1. This is the same discipline as writing golden cases, but applied to the environment instead of the agent.
Why can't the existing reward simply catch a bad actor? Because the failures that matter most for measurement barely move the reward. The paper counted 480 premature-disclosure failures and 134 goal-deviation failures. When a user gives up early, the agent's pass rate drops by about 0.52, so the reward notices. When a user over-shares, the pass rate changes by only 0.04 on average, and in some cells it goes up. The most common user failure is the one task reward is least able to see. A pass/fail score that only checks the ending cannot tell a solved task from a task the user half-solved for the agent.
Put numbers on one domain. Telecom has 114 tasks. With Sonnet-4-6 playing the user, the frozen GPT-5.5 agent scores a task reward of 0.947, which is about 108 of 114 tasks solved. With GPT-4.1-mini playing the user, the agent scores exactly the same 0.947: the same 108 wins. Now look at the actors. Sonnet-4-6 has a UFS of 0.474, so only about 54 of its 114 episodes kept fully to the script; GPT-4.1-mini has 0.737, about 84. Since at most 54 of Sonnet-4-6's 108 wins can be in-contract, at least 54 of the agent's 108 "wins" under Sonnet-4-6 came with a broken user script, against at least 24 under GPT-4.1-mini. Same scoreboard, two very different tests. (Counts are derived by multiplying the paper's reported rates by 114 and rounding.)
| User proxy | Cost per simulation | Agent task reward (mean of 4 domains) | UFS (mean of 4 domains) | On the cost–fidelity frontier? | Source |
|---|---|---|---|---|---|
| Sonnet-4-6 | $0.1841 | 0.784 | 0.778 | No: Gemini is cheaper and more faithful | paper |
| Gemini-3.5-Flash | $0.0558 | 0.796 | 0.864 | Yes, highest UFS | paper |
| Qwen3.6-35B | $0.0249 | 0.779 | 0.612 | No: third-highest reward, second-lowest UFS | paper |
| Gemma-31B | $0.0175 | 0.737 | 0.845 | Yes, cheapest proxy above UFS 0.84 | paper |
| Gemma-26B | $0.0147 | 0.720 | 0.759 | Yes | paper |
| GPT-5.4-mini | $0.0144 | 0.644 | 0.599 | No | paper |
| GPT-4.1-mini | $0.0063 | 0.720 | 0.752 | Yes, cheapest overall | paper |
The table shows why reward is a poor way to choose a simulator. Ranked by the agent's reward, Qwen3.6-35B looks like the third-best actor; ranked by fidelity, it is second-worst, because it over-shares. The paper therefore treats the choice as a cost–fidelity trade-off: pick the minimum fidelity your pipeline needs, then take the cheapest proxy that meets it. At a required UFS of 0.84 that is Gemma-31B at about $0.0175 per simulation; raising the bar to 0.86 moves you to Gemini-3.5-Flash at 3.2× the cost. For an RL run of 64 rollouts over 1,000 prompts, that is 64,000 simulated users per epoch, or about $1.1k versus $3.6k in user-side inference. Three of the seven proxies are strictly dominated in this setup: another tested proxy was both cheaper and more faithful at these prices.
The stakes are higher in training than in evaluation. In an exam, a lenient actor only inflates one grade. In multi-turn RL, the episodes are the training data, so an over-sharing simulated user can reward success without the information gathering the task was meant to teach. The paper names this risk but does not yet measure its effect on a trained policy. Among successful episodes, agents facing over-disclosing users made 1.06 fewer tool calls and asked 0.26 fewer questions, with identical reward. Across the 374 tasks all seven proxies completed, the agent's mean number of turns tracked UFS with a correlation of 0.955, against 0.594 for task reward. Across these seven proxies, a more faithful user went with more agent turns much more closely than a higher task reward did (this compares proxy-level averages, not single episodes). The same point applies to offline evals in production: if a simulated user drives your regression suite, the simulator is a versioned dependency that needs its own check.
The paper is honest about its limits, and they matter for how far to trust the numbers. The rubric writer and the grader are the same model (Claude Opus 4.8); the manual audit was done by one author on ten tasks plus 50 judgments; each task got one rollout per proxy; and serving prices change over time. UFS measures whether the actor followed this benchmark's script, not whether it behaves like a real person. A blinded human study and an RL experiment that trains against high- and low-fidelity users are listed as the next validation steps.
Goes deeper in: AI Agents → Evals & Diagnostics → The 4 Eval Failure Modes
Related explainers
- LLM-as-a-Verifier — A continuous verifier score — the other half of the measuring instrument: how the grading model turns a trajectory into a score.
- Beyond Static Leaderboards — Predictive validity — another case where an agent's aggregate score stops meaning what it appears to mean.