Ten advances in mathematics and theoretical computer science
Simon Willison Simon Willison ● Covered by 2 sources
OpenAI used an internal version of its next major model to find solutions to ten mathematical problems that had seen no progress for at least a decade, spending less than $2,000 per problem on token costs. The company published Lean 4 formalizations of the results, a paper describing the solutions, and an LLM-generated PDF reconstructing the proofs from reasoning traces. This demonstrates AI's potential to handle technical mathematical work at scale, aligning with mathematician Terence Tao's vision of "big mathematics" as a human-AI collaboration where machines handle technical tasks while humans focus on creative aspects.
Why it matters
Ten advances in mathematics and theoretical computer science A few days ago it was Anthropic discovering cryptographic weaknesses with Claude using Mythos Preview, spending $100,000 on tokens and with prompts that included "again we are not looking for low hanging fruit, we want proper research to find genuinly hard findings." Now it's OpenAI's turn to flex. They set "an internal version of Astra, our next major model" on finding solutions to ten mathematical problems that "have seen no progress on the main result for at least a decade". They claim to have spent less than $2,000 at GPT-5.6 Sol token prices on each one. (No news on how many problems they spent $2,000 on without reaching a solution though.) The openai/ten-proofs repository has Lean 4 formalizations of their results, and there's also a paper describing the solutions and an additional LLM-generated PDF where the model "reconstructs how the proof came together" based on the unpublished reasoning traces. That's a decent leve