TLDRocket
Sign in

Mistral NeMo

Mistral AI Covered by 2 sources

Mistral and NVIDIA just dropped Mistral NeMo, a 12B model with a 128k context window. It's Apache 2.0 licensed and swaps right into anything running Mistral 7B.

Based on reporting by Mistral AI — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Mistral AI teamed up with NVIDIA to ship Mistral NeMo, a 12-billion-parameter model that punches well above its weight class. The headline number is the 128k token context window, which lets it chew through entire codebases or lengthy documents without losing the thread. Mistral says its reasoning, coding, and world-knowledge scores beat both Gemma 2 9B and Llama 3 8B, and because it uses a standard architecture, teams can drop it straight into existing Mistral 7B setups with no retooling.

What sets this release apart from a routine model bump is the multilingual focus. Mistral NeMo was built to handle English, French, German, Spanish, Italian, Portuguese, Chinese, Japanese, Korean, Arabic, and Hindi with real competence, not just token support. That's a deliberate bet that the next wave of AI adoption happens outside English-speaking markets, and Mistral is backing it with a new tokenizer called Tekken.

Tekken, built on Tiktoken and trained across more than 100 languages, is the quiet star of this release. It compresses source code, Chinese, and several European languages about 30% more efficiently than the SentencePiece tokenizer Mistral used before. Korean compression doubles, Arabic triples. Against Meta's Llama 3 tokenizer, Tekken wins on roughly 85% of languages tested. Better compression means cheaper inference and longer effective context, so this isn't a cosmetic upgrade.

On the instruction side, Mistral ran an alignment pass that noticeably improves multi-turn conversation handling, precise instruction-following, and code generation compared to Mistral 7B, with GPT-4o used as the judge for evaluation. The model was also trained with quantization awareness baked in, so it runs in FP8 without losing accuracy, which matters a lot for anyone trying to keep inference costs down at scale.

Both the base and instruct checkpoints are on Hugging Face under Apache 2.0, and Mistral is also making the model available through its own platform as open-mistral-nemo-2407, plus as an NVIDIA NIM microservice on ai.nvidia.com. That's three different doors into the same model, which tells you Mistral wants adoption more than it wants control.

My take — AI-written commentary, not fact-checked reporting

Apache 2.0 on a genuinely competitive 12B model is the real story here, not the benchmark charts. Every time a lab this credible releases something this permissive, it chips away at the argument that frontier-adjacent capability requires a closed API and a credit card. The multilingual and tokenizer work is the unglamorous engineering that actually determines who gets to use AI cheaply outside San Francisco and London, and I'd like to see more coverage obsess over that instead of leaderboard positions.

Read more about this at: Mistral AI

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.