a16z Podcast

Mehtaab Sawhney & Mark Sellke on why AI reaches the results humans gave up on

Mehtaab Sawhney and Mark Sellke· Mathematicians at OpenAI at OpenAI
·~65 min·English·a16z
LLMReasoningAI Company
TL;DR

OpenAI's models are producing short, human-like proofs for open math problems, shifting the bottleneck from proving results to understanding and organizing them.

01Core Mental Model

The Renaissance of Reachable Results

Human mathematicians abandon uncertain approaches because their time is scarce; the model has no such cost, so it executes on the exact ideas people gave up on, opening a class of results that were always within reach but never worth the human gamble.

you're kind of gambling against the problem. You're like maybe I should try this approach but it seems really unlikely and just not worth my time.

Mehtaab Sawhney & Mark Sellke, a16z Podcast
Key Insight
The advance being described is not raw brilliance but a removed constraint: the model does not have to price its own time, so the risk-reward math that governs a human career simply does not apply to it.

02How It Reasons

It Reasons Like a Colleague, Not a Brute-Forcer

The released chain-of-thought shows the model pruning the search tree with judgment rather than trying everything, reading, in the mathematicians' words, shockingly like the notes of an expert human.

it's reasoning kind of shockingly like an expert human would.

Mehtaab Sawhney & Mark Sellke, a16z Podcast
Key Insight
That the reasoning is legible matters as much as that it works: a proof a human can read and check is one the field can build on, unlike an opaque machine search that lands on an answer no one can follow.

03Why It Reasons Better

A Polluted Context Traps the Human, Not the Model

A human who spends weeks on a wrong path gets stuck in it, because the failed attempt is wired into how they now see the problem, while the model can simply start a fresh session that never saw the dead ends.

as a human if you have some like wrong path you go down for a while it can be hard to like rewire your brain to like start over and like try a different path.

Mehtaab Sawhney & Mark Sellke, a16z Podcast
Key Insight
This reframes 'running parallel agents' as something humans structurally cannot do: you cannot un-know an approach that failed, but you can always spin up a model instance that never learned to distrust it.

04The Flagship Result

The Sphere-Packing Bound, Pinned Down Exactly

On the decades-old question of how densely spheres pack in high dimensions, the model did not solve packing outright but proved that a classic linear-programming bound is asymptotically the best that framework can give, an equality that explains a limit earlier work had only guessed from numerics and tightens the known bound on density.

I had actually thought about this problem for about six months at some point when I was a graduate student and yeah just I remember making like absolutely zero progress on it.

Mehtaab Sawhney & Mark Sellke, a16z Podcast
Key Insight
The value was not a new packing number but a reason: a limit previously supported only by numerics now has a short proof one mathematician called exactly the right approach; the exact density in high dimensions is still open, but the best bound this framework can reach is now known.

05Taste and the Harness

One Model for Taste, One for the Grind

Because today's models are task-oriented and stop when the task is done rather than when the mathematics is exhausted, the mathematicians floated splitting the work: one model to supply judgment and direction, another to grind out the long-horizon proof.

one model that's responsible for taste and one model that's responsible for going out and like you know working for a long time at solving a hard problem

Mehtaab Sawhney & Mark Sellke, a16z Podcast
Key Insight
The open question underneath is whether taste is a separate faculty or just a byproduct of getting good at tasks, and if it can be split off, a supervisor model that supplies direction is a concrete early shape for recursive self-improvement.

06A Surprise

Short and Elegant, Not Thousand-Page

The fear was that machine proofs would be unreadable thousand-page monsters; instead they have been short and elegant, with the non-sofic-group counterexample running about 15 pages where the related human disproof took 250 pages on top of 200 more.

you're kind of like afraid that they're going to like generate all these thousand page things and you're just like never going to be able to understand it. But it's been kind of the opposite.

Mehtaab Sawhney & Mark Sellke, a16z Podcast
Key Insight
One of the mathematicians put the inversion bluntly, that only humans can generate 200-page proofs right now, which suggests the models are finding the compressed core of an argument that people reach only through long, effortful detours.

07The Limit

A Ceiling That May Never Fall

Even granting exponential improvement, the hardest problems may stay out of reach; the mathematicians think it is plausible AI will never solve something like P versus NP, which could push the field toward its deep mysteries and away from the routine ones.

it's plausible we'll never solve something like P versus NP

Mehtaab Sawhney & Mark Sellke, a16z Podcast
Key Insight
A permanent ceiling would be oddly reassuring for the discipline: if the top of the difficulty range stays unreachable, mathematics does not get finished, but likely gets rebalanced toward the few problems that still resist any amount of compute.

08What Changes

The New Scarce Skill Is Understanding

When proving a result was the hard part, understanding it came along for the ride; now that models can produce the proof, the scarce and valuable human work becomes absorbing, organizing, and explaining the mathematics they generate.

of course models are going to help us produce exponentially more mathematics but they also make it much easier to absorb it.

Mehtaab Sawhney & Mark Sellke, a16z Podcast
Key Insight
There is a self-correcting loop here they point to: the same models that flood the field with new results are also the fastest way to digest them, so the tool that creates the understanding debt is also the one that helps pay it down.