The mathematical community is raising fundamental questions about how much trust can be placed in the 719 AI-generated mathematical manuscripts released by OpenAI. The questions arose after a recent paper pointed out that OpenAI’s previously announced Navier–Stokes proof does not exactly match the computer-verified Lean code.
Key points
A sign error came to light, reducing the number of OpenAI’s AI math manuscripts from 722 to 719.
Three British mathematicians found two places where the written proof and Lean code diverge in the Navier–Stokes proof.
The paper’s authors do not conclude that the proofs themselves are wrong, but emphasize that machine verification alone cannot replace peer review.
OpenAI’s AI math release
OpenAI released a large collection of mathematical proofs on October 6. An analysis published on Thursday subsequently pointed out that the work diverged in some respects from guidelines an independent panel of mathematicians had issued for AI labs. The catalog initially contained 722 papers, but OpenAI withdrew three after sign errors were found in them, invalidating one line of reasoning and, in turn, two additional papers that relied on it.
The guidelines were issued on September 29 by the nine-member Advisory Group on Mathematics and Artificial Intelligence. Their very first sentence urges labs not to “test advanced mathematical problems on closed commercial models.” Meanwhile, OpenAI’s repository says these papers resulted from applying closed internal models to undisclosed research problems.
Only 10 of the results are accompanied by explanations summarizing the model’s reasoning process. Just 42% of the final major results have been formalized in Lean, a programming language that allows mathematical proofs to be verified by machine.
Read also: How much does Anthropic’s CEO earn a year? Compensation package revealed in IPO filing
The Navier–Stokes proof: Where did it go wrong?
**Alexander Bastounis** of King’s College London and Fabian Circelli and **Anders Hansen** of the University of Cambridge closely examined proofs related to the Navier–Stokes equations, which OpenAI announced it had “solved” last September. They found two places where the written paper and the Lean code diverged. In one of them, the Lean code was shown to prove only a weaker form of the estimate than the paper claimed.
The authors refrained from judging whether the written proofs were right or wrong. The issue is narrower: Lean can verify that the final theorem holds, but it cannot guarantee that every step in the human-readable argument is valid. They emphasize that AI-assisted proofs of this kind must therefore still undergo human peer review.
Fields Medalist **Terence Tao** wrote on October 6 that current problems are being solved by “AI prompters,” who do not understand the results well enough to handle questions and answers or even give seminar talks properly. Advisory group member and Harvard professor Melanie Matchett Wood likewise said that by the time such results are made public, “nobody yet understands the whole thing, and the real work is only just beginning.”
OpenAI’s proof sparks a dispute over trust
The advisory group has made clear that the ultimate judge of how well OpenAI followed its recommendations is the broader mathematical community as a whole. This controversy over verification comes after a month in which tensions had already been running high.
OpenAI announced its Navier–Stokes result on September 8, prompting an immediate dispute over credit with **Tristan Buckmaster**, a mathematician at New York University. Three days after the announcement, 25 Fields Medalists also issued a joint statement warning that the mass production of AI-generated mathematical results could harm the academic ecosystem.
Next article: OpenAI’s 722 published AI math papers are now for mathematicians to evaluate
