TLDRocket
Sign in

Developing Post-Training Technology to Adapt the Largest Open Foundation Model to Country-Specific Specifications

Sakana AI

Sakana AI adapted top open AI models for Japan, calling the results Namazu, and launched a chat app to run them. One base model refused 72% of political questions; the tuned version answers almost all of them.

Based on reporting by Sakana AI — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Sakana AI has been quietly building post-training techniques meant to take the world's best open-weight foundation models and reshape them to fit a specific country's culture, values, and safety requirements. The first public result of that work is a prototype model series called Namazu, released in alpha, paired with a new consumer chat product called Sakana Chat that puts the models in front of actual users.

The backdrop matters here. Pretraining costs have climbed to the point where only a small cluster of players, mostly in the US and China, can realistically compete at the frontier. At the same time, more of those frontier models are being released as open weights. That combination pushes the interesting work downstream, into post-training, where a company can take someone else's expensive model and bend it toward a specific market's expectations without redoing the enormous upfront training run.

Sakana's argument is that foreign-built models inevitably carry the ideological leanings and information controls of wherever they were made. So the company built its own datasets targeting the cultural and social context of Japan, then ran a correction pass on top of several base models. The base models themselves were picked simply for being strong performers at the time, and Sakana says the technique isn't tied to any one of them, meaning future versions could sit on top of whatever model happens to be leading later.

On straightforward capability tests, things held steady. Namazu scored close to its base models across reasoning, knowledge, and coding benchmarks including AIME'25, MMLU-Redux, GPQA Diamond, LiveCodeBench, and IFEval, meaning the post-training didn't cost much in raw ability. The more striking numbers show up on Sakana's own neutrality and factual-accuracy benchmark, built around politically sensitive Japan-related history and diplomacy questions. Base DeepSeek-V3.1-Terminus refused to answer 72% of those questions outright. After Sakana's post-training, the resulting Namazu-DeepSeek-V3.1-Terminus refused almost none of them, while reportedly improving on both neutrality and factual coverage rather than just talking more.

The strongest of the Namazu models, that same DeepSeek-based version, was also run through Japanese-language benchmarks — Nejumi Leaderboard4, Swallow LLM LeaderBoard v2, and JamC-QA — and landed roughly in line with both its base model and similarly sized rivals. Sakana says a full technical report with detailed scores, plus open weights for several Namazu variants, is coming later.

Sakana Chat itself adds web search on top of the models and went through a beta test with around 1,000 users before this launch, feedback the company says shaped the final tuning. Training ran on GMO Internet's GPU cloud resources over October and November 2025, and Sakana credited the broader open-model work from DeepSeek, Meta, and OpenAI as the foundation this project builds on. The company frames this as a first step toward bigger ambitions: combining multiple models and agent techniques into something beyond a chat app.

My take — AI-written commentary, not fact-checked reporting

The number that actually matters here is the 72% refusal rate baked into the untouched base model — that's the real story about how much political filtering ships quietly inside frontier models before anyone calls it censorship. Betting on post-training rather than pretraining as the place to fight over sovereignty makes sense when only a handful of labs can afford to train from scratch, and Sakana at least says the technique isn't locked to one model. But swapping one country's blind spots for another's, even carefully, deserves more scrutiny than a alpha release with promised benchmarks 'coming later' is going to get.

Read more about this at: Sakana AI

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.