TLDRocket
Sign in

Cohere Releases North Small Translate: A 218B MoE Translation Model That Scores 83.6 on WMT26 Across 50 Languages

MarkTechPost Asif Razzaq Covered by 2 sources

Cohere just released North Small Translate, a 218B open-weight translation model for 50 languages. It’s free on the API for now, and Cohere says it beats Google Translate and DeepL on its own tests.

Based on reporting by MarkTechPost, Asif Razzaq — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Cohere has put out North Small Translate, an open-weight machine translation model built by Cohere and Cohere Labs. It’s a sparse Mixture-of-Experts system with 218B total parameters and 25B active at a time, aimed squarely at one job: translation across 50 languages, from Albanian to Vietnamese.

The headline number is Cohere’s own WMT26 result: 83.6 averaged across all languages. The company says that tops DeepL, Google Translate, and several open models including GLM 5.2 and Mistral Large 3. There’s a catch, of course. These are Cohere-reported scores, judged with GPT-5.6-Sol, so they’re a vendor benchmark until someone else runs the same gauntlet.

Still, the model does more than win a blog post. Cohere says it can be used free through its API until rate limits kick in. It can also be self-hosted for non-commercial use, or licensed commercially. For deployment, that matters more than a flashy chart. Translation is useful when it is boring, dependable, and easy to plug into a workflow.

The model also comes with some practical muscle. In Cohere’s tests, it produced 112 output tokens per second on one setup versus 81 for Gemma 4 31B, and 39 versus 30 at higher concurrency. On long documents, it scored 48.9 on a two-chapter translation task, ahead of Google Translate at 21.3 and Gemma 4 31B at 19.4. Cohere also says the model reaches 80.1 at a cost of $0.000676 per task, far below Gemini 3.1 Pro Preview (high), which it puts at $0.038928.

Under the hood, this is a decoder-only sparse MoE Transformer with 128 experts, 8 active per token, plus shared experts. It uses a mix of sliding-window and global attention, and Cohere says the attention layout first appeared in Command A. The model has 16K input and 16K output tokens of text-only context, and Cohere is also offering three checkpoints for self-hosting, including a 4-bit version that runs on 1x B200 or 2x H100.

My take — AI-written commentary, not fact-checked reporting

Cohere is doing the sensible thing here: stop worshipping general-purpose models and ship a translation model that knows its job. That’s the part the industry keeps forgetting while everyone chases one model to rule them all. Translation is infrastructure, not a demo reel, and sovereignty talk lands better when the product is actually deployable.

Read more about this at: MarkTechPost

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.