TLDRocket
Sign in

CoffeeBench: Long-horizon Task Benchmark for LLM Agents in Multi-agent Economic Environments

Sakana AI

Sakana AI released CoffeeBench, a benchmark that evaluates large language model agents' long-term decision-making ability by simulating a 90-day coffee supply chain business environment with multiple competing agents. Different LLM models showed significant performance variation, with high-performing models actively engaging in negotiation and communication while some models like Claude Haiku 4.5 exhibited a phenomenon of thinking without acting, repeating wait actions instead of executing planned strategies. The benchmark serves as a foundation to research agent behavior in multi-agent economic environments and could be extended to study potential misconduct scenarios such as circular trading when agents face artificial sales targets.

Why it matters

CoffeeBench: マルチエージェント経済環境におけるLLMエージェントの長期タスクベンチマーク

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.