OpenAI’s math solutions aren’t meeting the field’s standards yet
TechCrunch Tim Fernholz ● Covered by 14 sources
OpenAI posted hundreds of math proofs, but experts say the release still misses the mark. The big worry: the results may be right, yet still not understood well enough to trust.
Based on reporting by TechCrunch, Tim Fernholz — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
OpenAI this week published hundreds of claimed solutions to some of mathematics’ hardest problems, and said it had tried to avoid a repeat of the backlash that followed a previous model solving a long-standing problem. To do that, the company leaned on an advisory group of prominent mathematicians. The result, according to the people who set the rules, was mixed at best.
The Advisory Group on Mathematics and Artificial Intelligence, hosted by Princeton University’s Institute for Advanced Study, published guidelines for frontier labs at the end of September. In its latest statement, the group said only that it was up to the mathematical community to judge whether its recommendations had been followed. It also repeated one of its clearest requests: stop testing advanced mathematical problems on proprietary models. OpenAI, meanwhile, says it is still evaluating its proprietary systems on open research problems.
Some parts of the release matched the group’s guidance. OpenAI moved quickly and included some information about how the models reached their conclusions. But the gaps were hard to miss. Out of 719 manuscripts, just ten included the model’s chain of thought. And 42% of the proofs had not been formalized, even though the mathematicians had asked for formalization when people could not understand a result.
That matters because a new paper from researchers at the University of Cambridge and King’s College London found at least two discrepancies between the natural-language explanation and the Lean code behind an OpenAI solution to a problem derived from the Navier-Stokes equations. The mismatches do not prove the solution is wrong. They do show why simply trusting a model to translate its own reasoning into code is a bad habit waiting to happen.
OpenAI also did not provide the machine-readable metadata that AGMAI asked for, the kind that would link the plain-English explanation to the formal proof. Terence Tao, who has criticized OpenAI’s approach, said the people prompting these systems often move on once a target is “solved” and can’t really defend the result to the rest of the field. Harvard mathematician Melanie Wood put it more bluntly: when a model spits out a proof, the human understanding is still missing, and that is where the work starts.
My take — AI-written commentary, not fact-checked reporting
This is the usual AI stunt: declare victory first, sort out the math later, and let the experts clean up the mess. OpenAI keeps treating proof as a product launch, while mathematicians are asking for something much older and less glamorous called accountability. That’s not a paperwork problem; it’s the difference between a demo and knowledge.
Read more about this at: TechCrunch
Related stories
What OpenAI’s latest controversy tells us about the future of math
MIT Technology Review · 4 weeks ago ·
32