Latent Space

Weinstein & Handa on Computer Use Beating the Average Human

Ari Weinstein & Nikunj Handa· Leads of Computer Use and the API platform at OpenAI at OpenAI
·~40 min·English·Latent Space
AgentsInferenceAI InfrastructureAI Company
TL;DR

At OpenAI's DevDay, the leads of computer-use agents and the API platform explain how agents that drive ordinary software passed the average human on most tasks, where the gains come from — better models, richer inputs, and harness engineering together — and how OpenAI is packaging all of it into low-level cloud primitives developers can build on.

01The Core Claim

Computer Use Already Outpaces the Average Person

<strong>Agents that drive software are now faster than the average person at most tasks</strong>, and the next frontier is matching expert users.

now computer use is like faster at accomplishing tasks than like the average human probably in in most cases.

— Ari Weinstein, Latent Space
Key Insight
The claim is bounded on purpose: faster than the average human at most tasks, not yet faster than experts. That gap is the roadmap. Once an agent matches an expert's speed, delegating a task stops being a tradeoff you weigh and becomes the default you reach for.

02Why It Matters

One Agent Can Use Every Piece of Software

<strong>Because all software was built for humans, an agent that can operate a computer can do anything a person can</strong> — no per-app API required.

it makes it so that the agent can do anything you as a as a person can do. Um, because all the software in the world was designed for humans and now agents can use that same software uh and you can delegate to the agent.

— Ari Weinstein, Latent Space
Key Insight
This is why computer use sidesteps the integration problem that limited earlier agents. Instead of waiting for every service to ship an API for a task, the agent uses the same graphical interface a person would — so its reach extends even to apps and features that expose no API at all, like YouTube's A/B testing or community posts.

03What Actually Changed

The Delta Is Recovery, Not Getting Started

<strong>A year ago models could start a task but got stuck when something broke; now they debug, retry, and introspect their way through.</strong>

I think the biggest delta that I see is before they could like reliably start tasks but then they would run into problems and now they're really good at debugging. They're really good at trying again introspecting what is and isn't working.

— Ari Weinstein, Latent Space
Key Insight
Real-world tasks are mostly recovery, not the happy path. A model that starts well but freezes at the first unexpected dialog or failed click is a demo; one that notices the failure, forms a new plan, and tries again is a product. The jump in reliability came from closing that recovery gap, not from faster first steps.

04The Engineering Trick

Appshots Carry What a Screenshot Drops

<strong>A screenshot drops the information an agent needs — a link's target, a truncated title — so an appshot also hands the model the accessibility tree</strong>, the same structured data built for human screen readers.

if you take a screenshot of a web page that has a link, the screenshot doesn't include where the link goes. It doesn't include, you know, maybe you take a screenshot of your calendar, the event titles are truncated

— Ari Weinstein, Latent Space
Key Insight
The unlock is representation, not vision. Pixels throw away structure the model has to re-derive; the accessibility tree hands it the links, labels, and full text directly, and more token-efficiently. The model still mixes modalities — screenshots, accessibility, even Playwright — depending on the task, but richer structure plus letting it write code to act on a whole page at once is what replaced the slow screenshot-scroll-screenshot loop of earlier computer use.

05The Favorite Use Case

The Agent Becomes Its Own QA

<strong>With computer use, the agent that builds the software can also run it and test it</strong> — so you stop being the QA step in your own coding loop.

One of my favorite use cases for computer use actually and one that we see a lot in the wild is computer use letting the agent actually test the software that the agent has built which is far more consequential than it sounds

— Ari Weinstein, Latent Space
Key Insight
This closes the loop that still trapped coding agents: they could write code but not confirm it worked, leaving a human to click through and report bugs. An agent that can actually operate the running app turns build-test-fix into a cycle it can run on its own — so work can reach you already tested, as it does for the team dogfooding it.

06The Speed Play

The Decisions Model: Subtract, Don't Retrain

<strong>OpenAI's fast 'decisions' API is not a new model — it's the existing Luna weights with reasoning turned off, outputs constrained, and questions batched in parallel.</strong>

we haven't trained like a new model for this. We're like building this purely on top of the same Luna weights that we have.

— Nikunj Handa, Latent Space
Key Insight
The lesson for builders is that a lot of speed is available without touching the weights. Constrain the output, drop reasoning you don't need, and run independent questions as a batch, and a general model becomes a fast classifier. The harder, still-open part is calibration — knowing how confident the fast answer actually is.

07The New Bottleneck

Once the Agent Is Fast, the World Is Slow

<strong>As computer use speeds up, more of a task's time is spent waiting on the environment rather than on the model</strong> — a non-trivial share of benchmark time is just waiting for the website to load.

let's say you're automating a task on door dash.com like a lot of the time is actually waiting for door dash.com itself to load

— Ari Weinstein, Latent Space
Key Insight
This is the agent version of Amdahl's law: speed up the model and whatever you didn't speed up — page loads, network round-trips, a slow chat agent on the other end — takes a bigger share of the clock. Since a non-trivial slice of a task is already just waiting, part of the next round of wins is less in the model and more in shaving the delay between 'the page finished loading' and 'take the next action,' event-driven wherever the platform allows it.

08The Bigger Picture

Building an AWS for Agents

<strong>The platform team is shipping low-level primitives — websockets, async tool calls, mid-turn steering, caching, compaction — as the AI-native equivalents of cloud building blocks</strong> for developers to compose.

I I used to work at Stripe before this and uh at Stripe a lot of the game was like building these higher level primitives and products on top of like the core payments primitives.

— Nikunj Handa, Latent Space
Key Insight
The through-line of the API half is an open question Nick keeps returning to: what altitude to ship at. Ship primitives too raw and every developer rebuilds the same harness; ship them too packaged and you box in what people can build. He's still deciding how much to pre-assemble — which repeated harness patterns should become first-class API objects — versus leaving developers to compose the raw building blocks, the same tension that shaped early cloud computing.