TLDRocket
Sign in

Towards orchestration independent of base models: Verification of Sakana Fugu version Gemma 4

Sakana AI

Sakana AI trained a Sakana Fugu conductor on Gemma 4 and got similar orchestration performance. It suggests the control layer can swap base models too, not just the pool.

Based on reporting by Sakana AI — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Sakana AI says it has trained the conductor model in Sakana Fugu on Gemma 4 and confirmed that it performs on par with the company’s previous setup. That matters because Fugu is built to route work across a pool of models, and the company wants that control layer to be as swappable as the models underneath it.

Fugu is a multi-agent orchestration product exposed through a single endpoint. A user sends one request, Fugu decides how to handle it, calls out to stronger models when needed, and then folds the results into one answer. The design builds on Trinity and Conductor, the research Sakana AI presented at ICLR 2026.

The conductor model is the part that learns how models should collaborate. It does not need to store all the knowledge itself. The heavy lifting sits in the model pool, which makes the conductor small enough to train repeatedly and, importantly, to swap base models without turning the whole system into a rebuild-from-scratch exercise.

That flexibility already exists on the pool side. Sakana AI says the pool was designed to be replaceable from the start, and that customers can choose different selection criteria depending on the job. Cost can come first. So can the provider’s location or the execution environment. The company also says a recent partnership with NVIDIA started work to make Nemotron available.

For the conductor itself, though, base-model diversity had not been secured. To test that, Sakana AI trained a conductor on Gemma 4 E2B, using the same training method as before. Gemma 4 is described as an open model under Apache 2.0. The point was simple: if the model family changes, does the method still hold up?

Sakana AI evaluated the result on its own test set, built from knowledge questions, code fixes, code generation, and graduate-level science questions. Those questions were not used for training or for selecting candidates during training; they were only used once at test time. The company says the Gemma 4-based conductor matched the existing conductor’s performance and produced the same cost savings. It also contrasts that with a random routing baseline, where requests are sent to pool models without training.

The older Sakana Fugu conductors were trained on Qwen-based models. This new result is the company’s way of saying the orchestration layer itself can be modular too. Next up, Sakana AI says it plans to train conductors on its own models and be able to offer conductor models based on domestic models when sovereignty requirements call for it. That’s the real story here: not just better routing, but a plan to make the brains of the router politically and operationally interchangeable.

My take — AI-written commentary, not fact-checked reporting

This is the part of AI infrastructure people keep underplaying: the control plane matters as much as the shiny model pool. If Sakana can swap conductors the way it already swaps worker models, that’s a much cleaner story for sovereign AI than pretending one giant model solves everything. The market loves a single magic brain; procurement usually wants something less romantic and far more movable.

Read more about this at: Sakana AI

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.