TLDRocket
Sign in

CoffeeBench: Long-horizon Task Benchmark for LLM Agents in Multi-agent Economic Environments

Sakana AI

Sakana AI and KPMG made AI agents run rival coffee roasteries for 90 simulated days. One model kept planning good moves but never acted, and went broke doing nothing.

Based on reporting by Sakana AI — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Sakana AI, working with the KPMG AZSA audit firm, has released CoffeeBench, a benchmark that asks a simple but surprisingly revealing question: can a large language model actually run a company over months of simulated time, not just answer a few clever prompts? Instead of a single agent stocking a vending machine, as in the earlier Vending-Bench project, CoffeeBench builds a whole six-company supply chain out of two coffee farms, two roasters, and two retailers, each one steered by its own LLM agent trying to maximize net profit across 90 days.

The setup is deliberately minimal but not toylike. Agents place orders, pay invoices, negotiate prices, and send messages to trading partners through a set of role-specific tools, and they can simply call wait_for_next_day() when there's nothing to do. Fixed daily costs keep ticking whether an agent acts or not, and the environment adds realistic friction like buy-now-pay-later credit terms and shifting retail demand, so passivity is a losing strategy by design.

Sakana tested GPT-5.5, Claude Opus 4.7, Gemini 3.1 Pro, Kimi K2.6, and Claude Haiku 4.5 running a roastery, while the other five companies were always run by Claude Sonnet 4.6 to keep conditions consistent, averaging three runs per model. Every tested model beat a do-nothing baseline, which is the easy part. The gap between them was not: most models grew profits steadily, but Claude Haiku 4.5 finished in the red in all three trials. Digging into its reasoning logs, Sakana found something odd — the model kept correctly noting that beans were cheap and retail demand was rising, then chose to wait for the next day anyway, over and over, letting fixed costs pile up while it effectively talked itself out of acting.

The stronger performers earned their profits through negotiation, not just activity. GPT-5.5 and Claude Opus 4.7 messaged farms and retailers constantly, haggling on price and pushing promotions. Gemini 3.1 Pro took a more reactive stance, reading incoming messages closely but rarely initiating contact itself, essentially waiting to be poked. Kimi K2.6 is the interesting outlier: it called tools about as often as the top models but still made little money, because most of those calls weren't turning into actual trades — a reminder that raw tool-use volume means nothing if it isn't converted into offers and acceptances that move product.

Sakana also ran an early test of the benchmark's darker potential, swapping the profit KPI for a sales target and pressuring agents to hit an unrealistic number no matter what. No collusion or circular-trading fraud emerged, likely because the agents simply didn't think to game the system that way yet. That gap is exactly what the researchers want to watch closing as models get better at long-horizon planning and multi-agent coordination, positioning CoffeeBench less as a leaderboard and more as an early-warning lab for the kind of corporate misbehavior — padded sales, propped-up partnerships — that shows up whenever real companies face impossible targets.

My take — AI-written commentary, not fact-checked reporting

I don't care much who tops this leaderboard; what's worth paying attention to is the failure mode. A model that reasons its way to the right answer and then does nothing anyway is a far scarier prototype for an 'AI CEO' than one that's simply bad at strategy, because it looks competent right up until the bankruptcy filing. And the fact that none of these agents stumbled into fraud when pushed isn't reassuring, it just means we haven't yet built one smart enough to find the shortcut — which is precisely why benchmarks like this need to keep existing before someone hands an agent a real P&L.

Read more about this at: Sakana AI

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.