a16z

Daniel Litt on why proving theorems is not the same as understanding them

Daniel Litt· Professor of Mathematics at University of Toronto
·~63 min·English·a16z
ReasoningLLM
TL;DR

Mathematician Daniel Litt on what AI can and cannot do in math: it applies known techniques brilliantly and produces correct proofs, but seems weaker at the intuition and human understanding he sees as the real point, and today's incentives reward the artifact instead.

01Core Mental Model

Math Is Understanding, Not Papers

Litt's worldview starts here: mathematics aims not merely at proofs or papers but at understanding, and a correct proof without human understanding leaves that goal unmet.

The goal of mathematics is not to produce mathematics papers. It's to produce some kind of understanding. Maybe some of that understanding resides in model weights. To me, that's like pretty unsatisfying.

Daniel Litt, a16z
Key Insight
This reframes every AI-math headline. A model that emits a correct proof has produced an artifact, not the thing mathematicians actually want. It is why Litt can call the field's fastest results real and still find them beside the point: the understanding they were supposed to carry can be missing entirely.

02How It Reasons

The Proofs Have No Move 37

The surprise is how ordinary the reasoning looks. Not an alien, superhuman leap, but a recognizable human chain of thought, just tireless and unusually well read.

It's not like there's some you know move 37 or whatever. Um it's like a human mathematician doing math. It's like a human mathematician doing certain types of math.

Daniel Litt, a16z
Key Insight
The absence of a Move 37, AlphaGo's famously alien stroke, is itself the finding: in the results Litt inspected, the models followed a recognizable human path rather than an alien one. That helps explain both their strength on familiar techniques and why it matters to test where the resemblance stops.

03The Taxonomy

Brilliant at Applying, Weak at Inventing

The models are extraordinary at applying known techniques, grinding long computations, and combining ideas from papers an individual mathematician might not have read, while seeming weaker at intuition, big-picture judgment, and finding the right question.

they seem like weaker in things like intuition or like having some big picture point of view

Daniel Litt, a16z
Key Insight
The dividing line is not difficulty but kind. Anything shaped like apply a known method, very well plays to their strength; anything that requires deciding what is even worth asking plays to their gap. That gap is exactly the fuzzy, curiosity-driven work Litt says produces most mathematical progress.

04Where It Breaks

Refuting Is Cheap, Proving Needs a New Idea

A false conjecture can fall to one clever construction, which the models handle well; many true conjectures Litt studies appear to require serious new ideas, where models currently struggle.

I think one reason the models might not be useful for some of these things is like the conjectures are true.

Daniel Litt, a16z
Key Insight
It inverts the intuition that a true statement should be the easier target. Litt's favorite autonomous result was a surprising counterexample: the community expected the statement to hold, and a construction from an unexpected area refuted it. The true conjectures he cares about, he says, tend to sit inside large frameworks that a genuinely new idea, not a clever search, has to crack.

05A Useful Weakness

Being Bad at Grinding Made the Proof Better

Handed a lemma he could not force himself to grind out, Litt reformulated it, found a better statement, and the models proved that fast, a cleaner result than the brute-force proof would ever have been.

our inability to prove it like led to an improvement in the result.

Daniel Litt, a16z
Key Insight
A model would happily produce the ugly ten-page grind, correct and insight-free. Because a human could not stomach it, the problem got reframed and a conceptual explanation appeared. The uncomfortable implication for automation optimism: human friction was not overhead here, it was the source of the better mathematics.

06The Incentive Trap

Math's New Paper Mill Runs on a Slot Machine

Today's incentives reward paper count, so researchers re-prompt a model against old conjectures until a correct proof drops out, and preprint servers fill with results that may be correct but show no evidence of meaningful human engagement.

you can do that by playing the slot machine until um the model produces a hopefully correct proof of such a result.

Daniel Litt, a16z
Key Insight
The tell is mode collapse: the same proof of the same theorem shows up several times within days, because the models keep landing on one path. Litt's worry is not limited to wrong papers: even a correct proof builds neither understanding nor human expertise when there is no evidence anyone meaningfully engaged with it.

07Reading the Signal

Short Proofs Are a Ceiling, Not Elegance

Model proofs are short and clever not out of taste but because short proofs are the only ones anyone can currently verify; the long, grindy ones cannot be checked, so an 800-page AI proof is almost certainly wrong.

the reason it's producing short clever things is just like that's what we can check.

Daniel Litt, a16z
Key Insight
This turns a common compliment into a diagnosis. People praise the models for finding short, elegant proofs; Litt reads the shortness as the edge of the checkable. Push a harness to emit a 250- or 800-page proof and reliability drops, because length is exactly where neither the model nor a human can catch the error.

08The Human Stakes

Superhuman Models Still Need the Whole Pipeline

Even if the models become robustly superhuman, Litt argues the field still needs a broad human community: the frontier rests on millions learning to think mathematically, and handing everything to the model risks one mathematician cloned a thousand times, not a million different ones.

you need a an entire mathematical community to support a small group of people who are on the frontier

Daniel Litt, a16z
Key Insight
The argument is about design, not capability. Something being optimal does not mean a hand-it-all-over system will choose it, and the diverse, curiosity-driven exploration that historically drives progress is not guaranteed to survive if the model picks the directions. Keeping humans engaged keeps that option open, and, Litt warns, guards against quietly becoming worse thinkers as the AI ascends.