AI models flub these intelligence tests. Can you fare any better?
MIT Technology Review Grace Huckins
AI still flubs some logic and visual puzzles. The weird part: humans now beat models on traps they’ve seen before.
Based on reporting by MIT Technology Review, Grace Huckins — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Puzzles have been part of AI from the start. That’s not an accident. If a system can handle chess, Go, checkers, riddles, grids, and visual trickery, it looks a lot less like a autocomplete engine with a swagger problem and a lot more like something approaching general reasoning. The catch is that the same old puzzle boxes keep revealing where the gloss runs out.
The contrast is sharp. A Columbia University team reported in late 2024 that even the best models solved only 18% of the New York Times Connections puzzles. By early 2025, some models were getting them almost perfectly. That sounds like a clean win for AI. But the article’s point is the opposite: progress in one puzzle doesn’t mean models understand puzzles in the same way people do.
Spatial reasoning is still a mess for them. In mental-rotation tasks, where you have to decide whether two images show the same object from different angles, current language models fall badly behind humans even when they can read visual input. The same weakness shows up in 3D thinking more generally. Architects and mechanical engineers think in space; these models mostly fake it.
Memory can also hurt. A 2024 study by Google and the University of Illinois Urbana-Champaign found that when researchers tweaked a classic Knights and Knaves puzzle just a little, models often latched onto what they seemed to remember instead of noticing the change. SimpleBench appears to trigger a similar problem. The questions look like harder things models may have seen during training, and the models miss the twist while humans catch it.
Even 2D grids expose the gap. ARC-AGI remains one of the best-known puzzle benchmarks, and models do better when the grid is turned into text rather than shown as an image. Research also suggests that when they do answer correctly, they may rely on tangled rules that don’t generalize well. Humans, meanwhile, tend to use simpler visual ideas. And while models have improved a lot on ARC over the past year, some of these puzzles still stop them cold.
There’s also a category where humans trip over their own instincts and models sometimes look steadier. The article points to puzzle suites built around intuition, where people rush into the obvious answer and get burned. On the more procedural side, scale matters: Apple researchers found that models can handle easy Tower of Hanoi and river-crossing problems, but start failing once the number of disks or people reaches six or more. Similar trouble shows up in logic grids, according to work from the University of Washington, Stanford, and the Allen Institute for AI. The takeaway is not that models are bad at everything. It’s that their strengths are oddly specific, and their failure modes are still very human-looking when they’re not.
My take — AI-written commentary, not fact-checked reporting
This is why the hype cycle around “reasoning” keeps looking silly. A model can ace one puzzle set, then stumble on a tiny rewrite, which is a great reminder that memorization dressed as thought is still just memorization. The real story isn’t that AI is getting smarter every month; it’s that benchmarks keep catching it wearing different shoes.
Read more about this at: MIT Technology Review