Sakana AI super-powers AI reasoning using Japan’s own Sudoku Puzzles
Sakana AI ● Covered by 5 sources
Sakana AI built a new benchmark that tests AI reasoning using tricky modern Sudoku puzzles, not just the classic 9x9 grid. Turns out top AI models like ChatGPT o3 barely solve any of the hardest ones—showing reasoning still breaks down fast.
Based on reporting by Sakana AI — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Sakana AI, the Tokyo-based research outfit, has turned to a very Japanese pastime to expose a very modern AI problem. The company just released Sudoku-Bench, a benchmark built not from math olympiad questions or PhD-level trivia but from Sudoku puzzles — including the wild, rule-bending "Modern Sudoku" variants that have evolved far beyond the newspaper grids Nikoli popularized in Japan in the 1980s.
The pitch is simple: standard academic benchmarks are getting saturated. Models like OpenAI's o3 and DeepSeek's R1 now cruise through tests that used to separate the impressive from the ordinary. Sakana argues Sudoku, especially the handcrafted, rule-heavy modern kind, is a better proving ground because each puzzle demands its own bespoke logic. Unlike chess or Go, where rules never change, a lot of these puzzles force a solver to understand an entirely new constraint system before making a single move. That's a different kind of reasoning than pattern-matching a familiar game.
And the early results are humbling. Sakana tested several leading models, both open and closed source, by feeding them partially completed puzzles. Most couldn't place a single correct digit on average. Only o3 managed to solve any puzzles at all, and even that was on the easiest tier — Sakana is blunt that a small success rate there doesn't translate to meaningful progress on the benchmark as a whole, since difficulty ramps up steeply from there. The failure pattern is telling too: models often build a solution that looks fine for most of the grid, then hit a late contradiction and either present a broken answer or insist the puzzle itself is flawed. Human experts, by contrast, hunt patiently for the puzzle's built-in
My take — AI-written commentary, not fact-checked reporting
This is a smarter benchmark idea than another leaderboard flex, mostly because it targets the exact failure mode everyone's been quietly ignoring: models that fake competence right up until the last step, then blame the puzzle. Partnering with actual world-championship-level human solvers for training data, instead of scraping more internet text, is the more interesting story here — it admits that raw scale hasn't solved reasoning and that better, denser human demonstration data might matter more than another parameter count increase.
Read more about this at: Sakana AI