Mathematicians Say OpenAI's Claimed Math Breakthroughs Miss the Standards the Field Just Wrote
OpenAI's release of 719 claimed solutions to open math problems falls short of the standards the field's own advisory group published weeks earlier, particularly on formalization and human understanding.
The Advisory Group on Mathematics and Artificial Intelligence (AGMAI), hosted by Princeton's Institute for Advanced Studies and comprising nine researchers, released guidelines for frontier labs at the end of September. Its first request was "to stop testing advanced mathematical problems on proprietary models." OpenAI's release explicitly says it is evaluating its proprietary models using open research problems in mathematics. The lab followed some principles — releasing results promptly and including information about how models reached conclusions — but not others. Just 10 of the 719 manuscripts included releases of the model's chain of thought. For proofs people don't understand, AGMAI recommended formalization, yet 42% of the released proofs had not undergone that process.
A paper from mathematicians at the University of Cambridge and King's College London documents at least two discrepancies between the natural language proof and the Lean code behind OpenAI's solution to a problem derived from the Navier-Stokes equations. Models first produce a natural language explanation, then attempt to express it in Lean, which confirms accuracy by compiling the proof as code. The discrepancies don't necessarily disprove either solution, but they undermine the assumption that models can formalize their own outputs without human involvement. The authors conclude that autoformalised Lean proofs "should not prima facie be trusted without the same peer review process and scrutiny that other proofs are subjected to."
AGMAI also asked labs to "include machine-readable metadata correlating the natural language and formal artifacts," which OpenAI did not do. The group suggested OpenAI should help fund the work of human mathematicians required to make the solutions meaningful. Terence Tao wrote that problems are "being solved autonomously by AI prompters who have no interest in the broader field itself once their initial target is 'solved', and do not understand the AI output well enough to answer questions on the result, give talks, or otherwise interact with the rest of the field." Harvard's Melanie Wood told TechCrunch that when a model spits out a solution, "there is not human understanding of them at the point of release, and now the work begins."