The Mathocalypse
Shtetl-Optimized ● Covered by 15 sources
Opinion — commentary, not a factual news event.
OpenAI dropped 372 math proofs, including one for the Unique Games Conjecture. Researchers say the real shock is that humans can’t read most of them yet.
Based on reporting by Shtetl-Optimized — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
OpenAI just detonated a pile of 372 math results, and one of them appears to settle the Unique Games Conjecture, a problem some mathematicians have chased for years. The release was recommended by an advisory group that includes Timothy Gowers and Edward Witten, which gives the whole thing a little more gravity than a typical model demo. But the headline isn’t only that a machine found proofs. It’s that, in many cases, nobody knows how to check them quickly without the machine’s help.
That uncertainty is already coloring the reaction from people closest to the work. Dana Moshkovitz, a complexity theorist who had worked on the UGC for her whole career, described the writing as psychedelic, confusing, and so badly organized that it was impossible to read without AI assistance. Her account is useful because it points to the next job: not cheering, not doom-posting, but digesting. There are Lean certificates for some of the results, though not all of them, so the formal side is there in at least part of the pile. Understanding, though, is another matter.
The UGC is only one item in a very crowded list. OpenAI’s set also includes L=BPL, a break on integer multiplication below O(n log n) time, a positive answer to the Unitary Synthesis Problem, parity not being in QAC0, a nearly 4th-power separation between randomized and quantum query complexity for total Boolean functions, a superquadratic separation between sensitivity and block sensitivity, an area law for 2D gapped Hamiltonians, matrix multiplication in O(n9/4) time, and a randomized polytime algorithm for approximately counting perfect matchings. There’s also an uncomputability result for solving polynomial equations over the rational numbers. Any one of these would have been a serious year for a field. Here they arrive as part of one dump.
There’s another wrinkle. This wasn’t presented as the work of some giant bespoke system with 10,000 agents and mountains of compute. The source says it was a current internal OpenAI model, using about 3 hours of GPT-Pro-level compute on average per problem, after being tried on about 8,000 problems. That means the model solved roughly 5% of the longstanding open mathematical problems it was shown. It also means the bottleneck is shifting fast: from finding a proof, to figuring out what the proof says.
And that’s before you get to the politics of publication. Virginia Williams and Josh Alman posted a separate preprint one day earlier with new subquadratic results for 3SUM and all-pairs shortest paths, and Anthropic took a different route there by letting the humans write up and announce the work. So the field now has two models: OpenAI’s messy public proof dump, and Anthropic’s curated handoff. One hands everyone the keys and the mess. The other keeps the keys inside the building.
My take — AI-written commentary, not fact-checked reporting
This is what AI looks like when it stops being a chat toy and starts doing actual work: a flood of results, followed by a smaller flood of exhausted humans trying to decode the output. The neat myth that only the “real” deep problems remain untouched is already cracking, and the people insisting otherwise sound less like skeptics than tenants arguing the building is on fire because the wallpaper still looks fine. The real fight now is over who gets to own the translation layer between machine proofs and human understanding.
Read more about this at: Shtetl-Optimized