TLDRocket
Sign in

OpenAI Five Benchmark: Results

OpenAI

OpenAI's Dota 2 bot squad beat a team of pro-level human players 2-1 in a live best-of-three. Yeah, the AI actually won this time, in front of 100k live viewers.

Based on reporting by OpenAI — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

OpenAI Five didn't just show up yesterday, it won. In a best-of-three against five humans sitting comfortably in the 99.95th percentile of Dota 2 players, the bot team took the match, with a live audience watching and roughly 100,000 more tuned in on the stream at once.

The opposing lineup wasn't a random pickup group either. Blitz, Cap, Fogged, Merlini, and MoonMeander made up the squad, and four of those five have played Dota professionally. This is a game notorious for punishing anything less than deep, fluid coordination, five heroes, constant map awareness, item timing, team fights that can swing in three seconds flat. Beating a team like that isn't a scripted demo win, it's a genuine result against people who've spent years mastering the game's chaos.

What makes this stick in the memory is the setting. Not a closed lab test, not a curated highlight reel, but a live event where anything could go wrong in real time and everyone watching could see it happen. OpenAI has run public Dota showcases before, but this one lands differently because the opponents were near the ceiling of human skill, not just strong amateurs.

The scoreline, two games to one, also matters. A clean sweep would suggest the humans never had a shot. A split result says something closer to the truth: this was a real contest, close enough that the pros clearly made the bots work for it, and the AI still found a way through.

My take — AI-written commentary, not fact-checked reporting

I'll say the obvious thing nobody wants to hear anymore: beating pros at a game is not the same as reasoning, safety-relevant capability, or anything resembling general intelligence, and treating every game milestone as a step toward AGI is exactly the kind of hype-inflation that makes actual safety conversations harder. Cool result for reinforcement learning at scale, sure. Evidence of anything scarier or grander, not really.

Read more about this at: OpenAI

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.