Weinstein & Handa on Computer Use Beating the Average Human
At OpenAI's DevDay, the leads of computer-use agents and the API platform explain how agents that drive ordinary software passed the average human on most tasks, where the gains come from — better models, richer inputs, and harness engineering together — and how OpenAI is packaging all of it into low-level cloud primitives developers can build on.
Computer Use Already Outpaces the Average Person
<strong>Agents that drive software are now faster than the average person at most tasks</strong>, and the next frontier is matching expert users.
now computer use is like faster at accomplishing tasks than like the average human probably in in most cases.
One Agent Can Use Every Piece of Software
<strong>Because all software was built for humans, an agent that can operate a computer can do anything a person can</strong> — no per-app API required.
it makes it so that the agent can do anything you as a as a person can do. Um, because all the software in the world was designed for humans and now agents can use that same software uh and you can delegate to the agent.
The Delta Is Recovery, Not Getting Started
<strong>A year ago models could start a task but got stuck when something broke; now they debug, retry, and introspect their way through.</strong>
I think the biggest delta that I see is before they could like reliably start tasks but then they would run into problems and now they're really good at debugging. They're really good at trying again introspecting what is and isn't working.
Appshots Carry What a Screenshot Drops
<strong>A screenshot drops the information an agent needs — a link's target, a truncated title — so an appshot also hands the model the accessibility tree</strong>, the same structured data built for human screen readers.
if you take a screenshot of a web page that has a link, the screenshot doesn't include where the link goes. It doesn't include, you know, maybe you take a screenshot of your calendar, the event titles are truncated
The Agent Becomes Its Own QA
<strong>With computer use, the agent that builds the software can also run it and test it</strong> — so you stop being the QA step in your own coding loop.
One of my favorite use cases for computer use actually and one that we see a lot in the wild is computer use letting the agent actually test the software that the agent has built which is far more consequential than it sounds
The Decisions Model: Subtract, Don't Retrain
<strong>OpenAI's fast 'decisions' API is not a new model — it's the existing Luna weights with reasoning turned off, outputs constrained, and questions batched in parallel.</strong>
we haven't trained like a new model for this. We're like building this purely on top of the same Luna weights that we have.
Once the Agent Is Fast, the World Is Slow
<strong>As computer use speeds up, more of a task's time is spent waiting on the environment rather than on the model</strong> — a non-trivial share of benchmark time is just waiting for the website to load.
let's say you're automating a task on door dash.com like a lot of the time is actually waiting for door dash.com itself to load
Building an AWS for Agents
<strong>The platform team is shipping low-level primitives — websockets, async tool calls, mid-turn steering, caching, compaction — as the AI-native equivalents of cloud building blocks</strong> for developers to compose.
I I used to work at Stripe before this and uh at Stripe a lot of the game was like building these higher level primitives and products on top of like the core payments primitives.