Daniel Litt on why proving theorems is not the same as understanding them
Mathematician Daniel Litt on what AI can and cannot do in math: it applies known techniques brilliantly and produces correct proofs, but seems weaker at the intuition and human understanding he sees as the real point, and today's incentives reward the artifact instead.
Math Is Understanding, Not Papers
Litt's worldview starts here: mathematics aims not merely at proofs or papers but at understanding, and a correct proof without human understanding leaves that goal unmet.
The goal of mathematics is not to produce mathematics papers. It's to produce some kind of understanding. Maybe some of that understanding resides in model weights. To me, that's like pretty unsatisfying.
The Proofs Have No Move 37
The surprise is how ordinary the reasoning looks. Not an alien, superhuman leap, but a recognizable human chain of thought, just tireless and unusually well read.
It's not like there's some you know move 37 or whatever. Um it's like a human mathematician doing math. It's like a human mathematician doing certain types of math.
Brilliant at Applying, Weak at Inventing
The models are extraordinary at applying known techniques, grinding long computations, and combining ideas from papers an individual mathematician might not have read, while seeming weaker at intuition, big-picture judgment, and finding the right question.
they seem like weaker in things like intuition or like having some big picture point of view
Refuting Is Cheap, Proving Needs a New Idea
A false conjecture can fall to one clever construction, which the models handle well; many true conjectures Litt studies appear to require serious new ideas, where models currently struggle.
I think one reason the models might not be useful for some of these things is like the conjectures are true.
Being Bad at Grinding Made the Proof Better
Handed a lemma he could not force himself to grind out, Litt reformulated it, found a better statement, and the models proved that fast, a cleaner result than the brute-force proof would ever have been.
our inability to prove it like led to an improvement in the result.
Math's New Paper Mill Runs on a Slot Machine
Today's incentives reward paper count, so researchers re-prompt a model against old conjectures until a correct proof drops out, and preprint servers fill with results that may be correct but show no evidence of meaningful human engagement.
you can do that by playing the slot machine until um the model produces a hopefully correct proof of such a result.
Short Proofs Are a Ceiling, Not Elegance
Model proofs are short and clever not out of taste but because short proofs are the only ones anyone can currently verify; the long, grindy ones cannot be checked, so an 800-page AI proof is almost certainly wrong.
the reason it's producing short clever things is just like that's what we can check.
Superhuman Models Still Need the Whole Pipeline
Even if the models become robustly superhuman, Litt argues the field still needs a broad human community: the frontier rests on millions learning to think mathematically, and handing everything to the model risks one mathematician cloned a thousand times, not a million different ones.
you need a an entire mathematical community to support a small group of people who are on the frontier