Agent·

Qualify agents by reliability, human review, and cost with READY — Minimum-cost oversight policy — What does it mean?

The news. On September 2, 2026, researchers posted READY or Not: Reliable Enterprise Agent Deployment to arXiv. READY is an open testbed that stops asking how often an agent finishes a job by itself and starts asking a deployment question: given a reliability level you must hit, what is the cheapest human-oversight policy that gets you there, and can that policy be certified on cases the agent has never seen? An end-to-end clinical-audit case study ran 16 agent systems across 750 cases through it. Read the paper →

Stay at the security lane for a moment. Both scanners came off the same certification bench with almost the same number on the card — 72.8 and 72.5 out of 100. On paper you would buy either one. But the bench is not the terminal. On the bench a machine is fed test bags one at a time with nobody standing behind it; in the terminal bags arrive all day, and every bag the scanner cannot settle goes to a table where a person opens it by hand.

What separates the two lanes is not how often the scanner is right — it is how often it is unsure. A machine that is confidently wrong sends few bags to the table and quietly lets things through. A machine that flags everything it cannot resolve sends a long queue to the table and clears the safety bar, but you pay for that queue in staff hours, every shift, forever. Two lanes can post the same bench score and still need very different tables behind them.

READY makes the table the unit of measurement. You hand it three things: the agent, the workflow with its own definition of a job done right, and a class of candidate oversight policies — the different rules you could use to decide which cases a human sees. It measures the reliability and the operating cost of the whole human-plus-agent system under each policy, then selects the cheapest policy that still clears your reliability target, and proves that choice statistically on cases the agent has never seen. What comes out is not a leaderboard row. It is a deployment profile: the reliability you get, how often a human must step in, and what stepping in costs. That is the shape of an answer you can take into an eval-driven rollout decision, and it is the same quantity a fleet already tracks as a service-level objective.

The clinical-audit study is where the argument stops being philosophical. Sixteen agent systems went through 750 cases, and the two systems a benchmark would have called a tie turned out to need very different tables:

SystemAutonomous accuracyHuman review to qualifyReliability target
System A (READY study)72.8%39.2%76%
System B (READY study)72.5%29.6%76%
Gap0.3 points (derived)9.6 points (derived)

Put a workload behind those percentages. Say the queue is 1,000 clinical-audit cases a month — an illustrative volume, held fixed so only the review rate moves. System A qualifies at the 76% reliability target only when 39.2% of cases go to a person, so a reviewer opens 392 cases. System B clears the same bar at 29.6%, or 296 cases. That is 96 extra cases a month for the same reliability on the same work. At an illustrative 20 cases an hour, System A costs 4.8 more reviewer-hours every month — bought with a 0.3-point accuracy advantage that the benchmark told you to prefer. Rank on the benchmark and you pick the system that costs more to run. The oversight queue is a standing line item that never appears in an agent's cost profile when you only count tokens and latency, and it is the axis a cost-versus-reliability plot is actually drawn against. (The 1,000-case volume and the 20-cases-per-hour rate are illustrative; the accuracy, review-rate and reliability figures are the paper's.)

The catch worth naming: READY reports an operating point, not a law. The 39.2% and 29.6% figures hold under the oversight policy the study evaluated, on one clinical-audit workflow, and a different policy class or a different workflow moves them. What transfers is the procedure — search the policies, price each one, keep the cheapest that clears the bar, then prove it on cases you held back. What transfers is the procedure, not the two percentages. Run that before the rollout rather than after, the same way shadow mode earns a rollout the right to happen, and the review rate stops being something you discover from an on-call rotation.

Goes deeper in: Agent Engineering → Production Evals → Eval-Driven Rollout

Related explainers

Frequently Asked Questions

Check what you knowMap your AI & GPU knowledge across every track — free, role-based