đź”® Why one AI is better than four #598
Exponential View Azeem Azhar
Opinion — commentary, not a factual news event.
Four AI agents argued together and often picked the wrong answer. One agent with all the facts almost always got it right, which is awkward for the “teamwork” pitch.
Based on reporting by Exponential View, Azeem Azhar — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
A new Anthropic test puts a familiar human problem in front of four AI agents: the group hears some shared evidence that points the wrong way, while only one or a few agents have the facts that would get them to the right answer. That setup is the classic hidden-profile problem, and it’s a nasty one for systems that are supposed to improve when they talk to each other.
They usually didn’t. Across most model families, the right answer showed up in only 17% to 36% of runs once the agents had to negotiate as a group. Give a single agent the whole evidence set, though, and it got the decision right almost every time. One exception stood out: Mythos 5 did better than the rest, landing at about 85%, for reasons the source doesn’t explain.
The pattern points to two separate weaknesses. First, these models are often too similar. The source says 30 agents on the same coding task can produce 18 identical git branch names, which is a pretty blunt way of showing low variance. Second, they don’t have the social machinery that helps human groups survive bad consensus: reputation, recourse, and some protection for the person who says the obvious thing is wrong.
That doesn’t mean the problem is solved. It means the current setup still rewards one strong agent more than a committee of near-clones. Thinking Machines’ answer is to keep different AIs around, built in different places with different values and purposes, to “keep the weirdness alive.” That sounds less like a slogan than a design requirement.
The other angle here is economics. In the State of AI report, a 10% token price cut was linked to a 12% to 18% rise in usage, which is real but not explosive. Patrick Saner’s point lands harder: the key isn’t the price of a token, it’s the cost of finishing a useful unit of work. That’s exactly where firms get awkward, because knowledge work is usually sold in bundles, not neat little tasks.
My take — AI-written commentary, not fact-checked reporting
This is why “more agents” is often just a fancier way to get more confident nonsense. Open models and closed models alike keep running into the same boring truth: if everyone thinks the same, the committee is theatre. The real moat is diversity, not a bigger group chat with better branding.
Read more about this at: Exponential View
Related stories
AI agents are agreeing and acting: machines are now smarter than humans. Their principals merely agree
Fortune ·
6
đź”® The curious economics of a $6 AI agent #597
Exponential View · 1 month ago ·
11