TLDRocket
Sign in

Reinforcement Learning

75 summarised stories about Reinforcement Learning, each linking back to the original source. Browse all topics →

+ Follow this topic

Friday, 17 July 2026

Learning to Orchestrate Agents in Natural Language with the Conductor

Sakana AI 12

Sakana AI trained a 7-billion-parameter Conductor model using reinforcement learning to manage and coordinate a team of other AI models by writing natural language instructions tailored to each task. The Conductor achieved 83.9% on LiveCodeBench and 87.5% on GPQA-Diamond, surpassing individual models in its pool while dynamically adapting its approach—using single queries for simple questions and constructing multi-step workflows for complex problems. This approach enables AI systems to leverage collective intelligence by learning to delegate tasks across diverse models rather than relying on fixed human-designed workflows.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.