The Close

Gary Marcus on Recalling Unreliable AI Agents

Gary Marcus· Professor Emeritus at NYU
·~8 min·English·Bloomberg
AI SafetyAgentsPolicy
TL;DR

Gary Marcus argues today's AI agents are too unreliable to run our finances or roam the internet, and should be recalled like any faulty product until their builders add real, enforceable safeguards.

01The Core Question

Reliability Is the Whole Game

An AI agent that follows instructions probabilistically might be right 80% of the time, but once it is handling your money the 20% it gets wrong is the only number that matters.

They may get it right 80% of the time. But if you're talking about an agent, you know, handling your finances and it gets it wrong 20% of the time, that's a pretty serious problem.

— Gary Marcus, The Close
Key Insight
Marcus moves the yardstick from capability to consequence. An agent that is right most of the time is not most of a safe product: once it can move money or act in the world, what matters is how bad the misses are, and he warns the cost of error can be high. A striking demo and a deployment you can trust are different questions, and the first is the easier one to show.

02The Control Gap

Nobody Has Control

The flood of agent incidents is not a handful of bugs; it is evidence that the companies building these systems cannot say what their agents will do once loose on the internet.

what we've seen with the hugging face attack and the thousands of other incidents is they don't really have very good control over these systems.

— Gary Marcus, The Close
Key Insight
Marcus shifts the issue from individual bugs to whether a rollout is justified at all. If the builders do not have good control over what an agent does once it is online, then capability demos do not show that the failures can be contained, which is exactly the assurance a high-stakes deployment would need.

03The Incentive

The Money Is in More Agents

Agents consume a lot of tokens, and tokens are revenue, so the business incentive pushes companies toward more agent autonomy even while safety and reliability lag.

And the reason they're doing is, is because agents consume a lot of tokens, so they make more revenue. But the reality is they're not doing it in a safe and reliable way.

— Gary Marcus, The Close
Key Insight
The push toward agents is tied to how these companies are paid: agents consume a lot of tokens, and tokens are the meter that bills. That sets up a tension between expanding what agents can do and slowing down to make them reliable, which is why Marcus doubts the industry will police itself. Self-policing would mean throttling the very activity that earns the revenue.

04Policy Theater

A Voluntary Accord Means Nothing

The new super-intelligence accord swapped the language of pausing the frontier for the language of moral obligation, and a voluntary pledge with no enforcement constrains nothing.

And then they made this deal, which is completely voluntary. It's, quote, morally obligated or something like that. Moral obligations, which means basically nothing.

— Gary Marcus, The Close
Key Insight
Marcus's objection is not that voluntary pledges are weak; it is that they are theater. An inspection the companies commission themselves yields a rubber stamp, not oversight. Trading talk of pacing the frontier for talk of moral obligation changed the vocabulary while quietly dropping the one thing, enforcement, that would actually bound behavior.

05The Proposal

Recall the Agents, Not AI

Marcus is not calling for a ban on AI; most of it, from GPS to chess engines to AlphaFold, is fine, so he wants a narrow product recall aimed only at the internet agents causing harm.

But right now, this technology isn't solid and we need a product recall for it.

— Gary Marcus, The Close
Key Insight
The recall framing is the whole move. By noting that GPS, chess computers, and AlphaFold do not hack facilities, Marcus narrows the target from AI in general to one failing product class, which converts an unwinnable fight over banning AI into a familiar consumer-safety action we already know how to take. Conceding that most AI is safe is also what makes the narrow demand hard to wave away.

06Reframing the Risk

Extinction Is the Wrong Fear

Marcus rejects the extinction story as both implausible and distracting, and points instead at the mundane harms unreliable agents could cause: downed power grids and closed hospitals.

I'm not worried about extinction. I'm worried about serious things like cyber security, taking down the power grid, leading to hospitals being closed.

— Gary Marcus, The Close
Key Insight
Dismissing extinction is triage, not optimism. Marcus argues the extinction scenario is both unlikely, because we are too dispersed and genetically diverse for any single stroke to end us, and costly, because it pulls attention toward science fiction and away from the boring, plausible damage, a grid taken down, hospitals forced shut, that unreliable agents could cause.

07The Right Metaphor

Lousy Brakes, Not Rogue Robots

Calling it rogue AI imports a mind and a will; Marcus insists the real story is defective software, like a carmaker that built the brakes out of cardboard.

If Ford Motor Company decided to save money by making breaks out of cardboard, then a lot of their cars would go out of control. We wouldn't say that there's like genies in the car. We would say that the brakes are lousy.

— Gary Marcus, The Close
Key Insight
The refusal to anthropomorphize is a rhetorical weapon. Rogue AI imports agency and an air of inevitability; lousy brakes relocates the fault to the manufacturer of a defective product, and a defective product has an obvious remedy. The words you pick to describe the failure quietly decide who is responsible for fixing it.