Can AI Learn Mathematical Intuition?
a16z
Setting the Stage on AI and Math 0:00
The conversation opens with the idea that mathematics is not really about producing papers, but about producing understanding, and that some kinds of understanding living only inside a model's weights would be an unsatisfying outcome. The guest, Daniel, is introduced as a professor of mathematics at the University of Toronto who has been outspoken about how his views on AI in mathematics keep shifting. The host frames the discussion around a key question: what has been the most impressive AI result in mathematics so far, and how should the mathematics community respond to this progress.
The Irish Unit Distance Problem 2:31
Daniel names his favorite fully autonomous AI result as the solution to the Irish unit distance problem, announced in mid-May. What made it stand out was its creativity: the mathematical community had believed a certain statement was true, but a counterexample was found, and the proof pulled in classical techniques from the 1960s that were new to this particular area of study, involving point configurations in the plane. Afterward, other mathematicians took those same imported ideas and used them to find counterexamples to other open questions, including the sum-product conjecture over the real numbers. Daniel says this is his test for judging whether a result is genuinely impressive: do the new ideas it introduces turn out to be useful for other problems afterward.
How Human Does the Reasoning Look 5:00
Daniel pushes back on the idea that the model's reasoning felt inhuman. He recalls that OpenAI released the chain of thought behind the unit distance result, and it read as recognizably human, similar to how he imagines his own thought process might look while solving a problem. Most AI-produced results he has studied read like something a human mathematician could have written, not like some alien or unexpected move. Where the models seem slightly inhuman is simply that they don't get tired and they know a great deal, not that their reasoning process itself is strange.
Comparing OpenAI and Anthropic 9:00
Daniel notes that OpenAI's and Anthropic's models seem fairly close in mathematical capability, with cases where one lab releases a solution to a problem and the other announces the same solution shortly after. He has used ChatGPT more than Claude, partly because ChatGPT became useful for research math earlier, with Claude catching up only around Opus 4.5 or 4.6. Both models, he says, are strong at applying known techniques, grinding through computation, and pulling together technical ideas from papers a human mathematician might not have read, but both are still weak on intuition and big-picture point of view. He also raises the point that the labs mostly scaled natural-language reasoning rather than training on formal, verified proofs in systems like Lean, which he suspects means these techniques will generalize well beyond mathematics.
Signs of Theory Building 11:00
Daniel describes trying to get models to do "theory building," a fuzzier, more conceptual activity than solving a stated problem. On their own, neither model is good at this, but given hints, they can produce something interesting. He notes that when he supplies hints it is hard to separate his own contribution from the model's, but suspects that whatever they can do today with a hundred bits of hinting, they might do unaided within six months. He sees this as a sign that models are starting to pick up some of mathematics' implicit, unwritten knowledge.
Problem Solvers Versus Theory Builders 14:00
Daniel lays out a rough taxonomy of mathematicians as problem solvers versus theory builders, placing himself closer to the problem-solver end. For him, an open problem is a kind of benchmark that measures a specific failure to understand something, and he gives the example of the Grothendieck period conjecture as measuring a gap in understanding differential equations. He describes his own work as often driven by analogies, such as the fruitful parallel between the homology of algebraic varieties and representations of fundamental groups, an analogy that has generated decades of mathematics by researchers like Carlos Simpson and Takuro Mochizuki. He also brings up the Birch and Swinnerton-Dyer conjecture, one of the Millennium Prize problems, as an example of a conjecture born from noticing a pattern in computed statistics on elliptic curves in the 1960s, calling it an early instance of data-driven mathematics.
Where AI Helps and Where It Doesn't 19:00
Daniel says AI is least useful on vague, long-running projects, the ones he has been thinking about for three, four, or five years, where there isn't yet a precise question to ask. For those, AI mostly functions as a substitute for Google, saving time on background reading rather than doing deep intellectual work. Where it has opened up new territory for him is coding: since he isn't a strong coder himself, having a tireless assistant has let him take on projects, especially ones needing massively parallel exploration, like testing a thousand examples at once, that he would have otherwise put off indefinitely.
Beauty, Science, and Personal Curiosity 21:00
Asked what drives his choice of problems, Daniel says he tries not to be motivated by beauty or aesthetics, since chasing elegance can limit a mathematician who assumes an ugly-looking proof must be wrong. He prefers to think of his work as doing physics with concepts, aiming for what will open up further understanding rather than what looks pleasing. He adds that mathematical progress broadly comes from many people pursuing their own personal curiosity, with the frontier of knowledge expanding as new ideas cascade into unexpected areas. He closes this thread by noting that models are often not useful on his own long-running problems partly because those conjectures tend to be true, whereas cases like the unit distance problem, where the belief turned out to be false, hinge on finding a specific counterexample construction.
Why new theory is hard for AI 24:30
The guest explains that models are good at finding constructions, such as counterexamples, but struggle when a problem requires building an entirely new theory or technique rather than porting over existing methods from other fields. He points to the unit distance problem as a case where the model pulled in an idea from an unexpected area, which made it feel like a genuinely creative result rather than a routine application of known tools.
Curricula for humans and machines 27:00
He argues that developing a theory or understanding a poorly understood object is a fuzzier skill than solving a well posed conjecture, making it harder to design rewards for it. Still, he expects continued growth in this area, partly because mathematics offers a huge range of conjectures of varying difficulty that can act as a kind of curriculum, similar to how human PhD students build skills through their advisors' feedback.
A story about a stubborn lemma 29:31
He recounts working with Gemini Deep Think on a paper where the model could not prove a needed lemma, so he worked out examples by hand and found a better, more elegant statement of it, which the model then proved quickly. He notes that a brute force ten page proof would have worked just as well logically, but he could not bring himself to grind through it, and that reluctance led to a more conceptual and satisfying explanation of why the result was true.
Understanding versus producing papers 34:30
He states plainly that the goal of mathematics is not to produce papers but to produce understanding, and that understanding sitting only in a model's weights feels unsatisfying to him. He stresses that a small group of frontier researchers depends on a huge pipeline of people learning to think mathematically, and that current incentives, which reward posting many AI assisted proofs of existing conjectures, do not encourage anyone to build that human capital.
The slot machine problem in research 36:01
He describes an experiment where he asked Codex to find and prove five recent conjectures in algebraic geometry, producing three correct but low quality papers in an hour, and notes that many similar papers now appear on arXiv, sometimes with three or more papers proving the same theorem within days because people are repeatedly running the same kind of prompt. He worries this reflects mode collapse onto familiar reasoning paths, and that subordinating mathematical exploration to what models happen to pursue could narrow the diversity that normally drives progress, since human mathematicians bring wildly different intuitions that models mostly draw from the same existing literature.
Keeping humans central to research 41:00
He argues that even if models become robustly superhuman at math, humans should still stay involved, because there is no guarantee that leaving research entirely to models produces the diverse, broad based exploration that has historically driven mathematics forward. He says the way to protect this is to keep a community with wide interests actively steering and using the models, so that outcomes remain shaped by human curiosity rather than by whatever path is most direct for the model to take.
Education, incentives, and staying sharp 46:01
Both speakers discuss how math's close tie to education makes it a useful lens for watching AI's effect on learning more broadly, mentioning a bimodal split emerging in classrooms between students who use new tools to deepen understanding and those who let the tools do their homework. They warn that new technology often produces outputs that are cheaper but lower quality than before, so institutions need redesigned incentives to ensure models are used to deepen human understanding rather than replace the effort of thinking altogether.
Slop and enthusiasm for math 48:30
Daniel notes that some of the low quality mathematical content flooding the internet comes from professionals facing pressure to publish, not just amateurs. He does not see the rise of amateur enthusiasm as a net negative. Even though it means more documents to sift through when checking whether a problem has already been solved, he sees more people getting excited about mathematics as a genuinely good thing.
The rank 30 elliptic curve result 49:30
The conversation turns to a new elliptic curve of rank 30, produced with AI assistance from a system called Fable, prompted by Levent Poge and a collaborator, with Ava Howell possibly involved as a third contributor. Daniel points out that the methods have not been disclosed, so it is hard to judge significance, and this kind of construction, historically pursued by mathematicians like Noam Elkies and later Aliz and Klagsbrun with a rank 29 example, is a fun but not top-tier result. He stresses that any mathematical result, human or AI-generated, can really only be judged in retrospect, once its difficulty and importance become clear.
Why AI proofs stay short 52:00
Daniel explains that current models tend to produce short, clever proofs mainly because they cannot yet reliably verify long, complicated ones. Ask a model if its own short proof is correct and it will sometimes admit it is wrong, showing real but limited self-checking ability. He suspects labs like OpenAI and Anthropic have solved more problems internally than they've released, holding back longer proofs that resist formal verification in systems like Lean. He points to an 800-page AI-generated proof claiming to resolve singularities in positive characteristic, which he says almost certainly cannot be correct, since no human or model has actually checked it.
Harnesses, reliability, and checking proofs 56:30
Pushing a model toward long proofs generally requires a "harness," a scaffolding setup that elicits extended output, and Daniel argues this often lowers reliability because it pressures the model to produce something rather than something carefully verified. He compares this to how humans actually check long papers, not line by line but by testing the overall structure and stress-testing it against known special cases. He recalls a flawed paper he and other experts caught not by finding the specific error but by sensing the argument implied something too strong to be true; current models still cannot perform that kind of big-picture skepticism. He mentions using a small personal harness with tools like Codex, but says he prefers using AI to help himself understand mathematics rather than to do it autonomously.
Raising a child around math 59:30
Daniel shares that his three-year-old daughter can count to about thirty reliably, is starting to add, and has picked up the platonic solids from a toy icosahedron, with group theory joked about as the next step. He believes the core reasons to learn math, thinking clearly and understanding the world, will remain valuable even as AI grows more capable, much like the reasons to read broadly in the humanities. He hopes to pass that love of the subject on to her, while acknowledging that educational institutions may need real change to keep those values intact as AI reshapes the field.
AI-generated summary. It can be wrong or incomplete - check anything that matters against the original.

