Recursive Self-Improvement through Multi-Agent Self-Supervision
Sakana AI ● Covered by 3 sources
Sakana AI says a model got better by judging teams of its own copies. That raises the awkward question: if the evaluator is flawed, who catches it?
Based on reporting by Sakana AI — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Sakana AI is betting that a model can teach itself by working with itself. The company’s new method, Multi-Agent Self-Supervision, has one shared model spin up a team of virtual subagents, search for better ways to coordinate them, and then use its own judgments to pick the best setup. The winning team’s runs are fed back into training so the main model absorbs what the team found useful.
The pitch is aimed at a real problem. When AI is pushed toward open-ended research tasks, human reviewers may not be able to reliably tell which answers are good, or how to push the system further. Sakana’s answer is to lean into the machine side of the loop instead of pretending a person can always sit in the middle.
The company says it ran two cycles of this process with a 27B open-weights model on synthetic open-ended research tasks. After the second cycle, the system’s score per output token was 1.2-1.6x the base model’s level across four research benchmarks. That is a meaningful gain, and it came from training on the collective output of the model’s own multi-agent setup.
But the same setup is also the obvious danger. If the model is both the worker and the judge, shared blind spots can look like progress and get reinforced during training. Sakana is unusually direct about that risk, arguing that homogeneous recursive self-improvement needs close scrutiny if AI safety is supposed to keep up with AI capability. The message is less “look, self-improving AI is solved” and more “we’ve found a promising way to make the loop tighter, so now we need to make sure the loop doesn’t lie to itself.”
My take — AI-written commentary, not fact-checked reporting
This is the kind of result that should make both open-model fans and AI safety people slightly nervous, which is usually a sign the work is interesting. Models grading their own homework is efficient, but it also gives failure modes a comfortable chair and a long-term lease. The industry has spent years chasing bigger loops; now it needs better brakes.
Read more about this at: Sakana AI