TLDRocket
Sign in

Emergent tool use from multi-agent interaction

OpenAI Blog

Agents trained in a simulated hide-and-seek environment discovered six distinct strategies and counterstrategies through multi-agent interaction without explicit instruction. The agents progressed from simple behavior to increasingly complex tool use within a single environment. This demonstrates that competitive multi-agent training can generate emergent complexity without direct supervision or predetermined task design.

Why it matters

We’ve observed agents discovering progressively more complex tool use while playing a simple game of hide-and-seek. Through training in our new simulated hide-and-seek environment, agents build a series of six distinct strategies and counterstrategies, some of which we did not know our environment supported. The self-supervised emergent complexity in this simple environment further suggests that multi-agent co-adaptation may one day produce extremely complex and intelligent behavior.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.