TLDRocket
Sign in

Cheaper, Better, Faster, Stronger

Mistral AI

Mistral just dropped Mixtral 8x22B, a giant open-weight model under the free-for-anything Apache 2.0 license. It's bigger than its predecessors but cheaper to run than you'd expect, and it beats other open models at math, code, and reasoning.

Based on reporting by Mistral AI — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Mistral AI has a habit of releasing serious hardware right when you least expect it, and Mixtral 8x22B is no exception. It's a sparse Mixture-of-Experts model with 141 billion total parameters, but here's the trick: only 39 billion of those are active for any given token. That sparsity is the whole point. It means Mixtral 8x22B runs faster than a dense 70B model while, according to Mistral's own benchmarks, outperforming every other open-weight model on the market right now, restrictive license or not.

The specs read like a wishlist. Fluent across English, French, Italian, German and Spanish. Native function calling, paired with a constrained output mode on Mistral's la Plateforme, which matters a lot if you're trying to bolt this into an actual production stack rather than just chatting with it. A 64K token context window, big enough to chew through long documents and still pull out precise details instead of hallucinating them. And on math and coding benchmarks — HumanEval, MBPP, GSM8K — it leads the open-model pack outright. The instruct-tuned version pushes that further, hitting 90.8% on GSM8K maj@8 and 44.6% on Math maj@4, numbers that would have sounded absurd for an open model two years ago.

What's arguably more significant than the benchmarks is the license. Mistral is shipping this under Apache 2.0, the loosest open-source terms available, no usage restrictions, no asterisks. That's a deliberate contrast to labs that dangle 'open' models behind research-only clauses or commercial carve-outs. Mistral is betting that raw openness, plus a genuinely efficient architecture, is a better growth strategy than gatekeeping.

The multilingual numbers back up the European pitch too. Mixtral 8x22B clearly outpaces LLaMA 2 70B on HellaSwag, Arc Challenge, and MMLU across French, German, Spanish and Italian, which is exactly the kind of result you'd want from a French AI company trying to prove it isn't just chasing English-language leaderboard scores. And because the base model is fully available, it's also positioned as fertile ground for fine-tuning, meaning this release isn't really the end product so much as a starting point for a lot of derivative work that's about to show up.

My take — AI-written commentary, not fact-checked reporting

I'll say the obvious thing everyone in Brussels wants said: a European lab just handed the world a frontier-grade open model with zero licensing strings, while the big closed-model shops keep lecturing us about safety before locking their weights in a vault. Mistral's efficiency angle is real and underrated — sparse MoE architectures are quietly becoming the actual competitive edge, not just parameter counts. My only worry is that 'truly open' releases like this become the exception that proves the rule, since most labs still treat openness as a marketing costume rather than a policy.

Read more about this at: Mistral AI

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.