TLDRocket
Sign in

Can AI beat a goldfish at calling the World Cup?

Rest of World Indranil Ghosh

AI chatbots are predicting World Cup matches, and losing to a goldfish in Toronto. Swimbappé's hitting 80% accuracy while ChatGPT, Claude, and friends languish around 50-60%.

Based on reporting by Rest of World, Indranil Ghosh — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Five major AI models — ChatGPT, Claude, Gemini, Copilot, and Perplexity — took a swing at predicting every single match of this World Cup, all 104 of them, sealed in advance by a Dutch research outfit called Onder and tracked on a public scoreboard named ScoreGPT. The models had rankings, form, injury reports, everything a human pundit would want. Their accuracy has hovered between 50% and 60% per match. A goldfish named Swimbappé, who lives in a Toronto storefront tank and picks winners by swimming toward a flag, is running at roughly 80%.

The AI didn't embarrass itself everywhere. Back in June, four of the models correctly named Spain, France, England, and Argentina as semifinalists before a single match had been played — an impressively clean call. But then reality got messy. Those same tools picked Norway as this tournament's dark horse and Brazil as its biggest letdown, yet not one of them saw Norway actually eliminating Brazil in the round of 16. Germany's loss to Paraguay caught all five models flat-footed; they'd confidently predicted a comfortable German win. And when USA Today put Copilot on the spot for four group-stage matches in a single day, every one of them ended in a draw — a result the model hadn't flagged as likely in any of the four.

Swimbappé, to be fair, never claimed to understand offside rules or squad rotations. He's just swimming toward flags. That's sort of the point. Animal oracles have been a World Cup sideshow since Paul the Octopus went 8-for-8 in 2010, including the final, and spawned a whole lineage of imitators: a camel in the UAE, a deaf cat in Saint Petersburg, a hawk named Shawk in Dubai who's called more than a dozen matches this tournament with a solid hit rate. Thailand has tigers and hippos on the case. Moo Deng, the pygmy hippo who became an internet celebrity, picked her semifinal winners by choosing between watermelons marked with flags.

The symbolism here isn't subtle. Onder isn't just testing whether ChatGPT can read a spreadsheet of odds — it's testing whether large language models can actually do prediction under real uncertainty, where injuries, momentum, and dumb luck matter as much as historical data. So far the honest answer is: not particularly well, and not obviously better than a fish with no data at all.

My take — AI-written commentary, not fact-checked reporting

I've spent enough time watching AI models confidently misfire on things they were built to be good at — pattern recognition, probability — that a goldfish beating five LLMs at football predictions feels less like a novelty and more like a small, useful humiliation. The lesson isn't that AI is useless; it's that chaos-heavy, low-signal domains expose how much of these models' confidence is theater. My money's still on the fish.

Read more about this at: Rest of World

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.