TLDRocket
Sign in

Mistral AI Releases Mistral Large 4 (Le Chonk): A 1.05T Parameter Multimodal MoE Model

MarkTechPost Asif Razzaq ● Covered by 7 sources

Mistral just previewed Large 4, a 1.05T-parameter AI with image input and a 1M-token window. Its best showing is cybersecurity, where closed models reportedly balk.

Based on reporting by MarkTechPost, Asif Razzaq — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Mistral AI has put Mistral Large 4, nicknamed Le Chonk, into public preview. It’s a huge model on paper — 1.05 trillion total parameters, 49 billion active per token, native image input, and a 1 million token context window. But the company is also being unusually practical about how to use it: the API is live now, while the weights won’t arrive until the end of October, so self-hosting is still off the table.

The pricing is meant to make that scale feel less absurd. Mistral says the API costs $1.36 per 1 million input tokens and $4.18 per 1 million output tokens, with cached input at $0.14 per 1 million tokens. That matters because long-context agent loops are exactly where a 1 million token window starts to stop sounding like a brag and start sounding like an actual product feature.

Under the hood, ML4 is a granular mixture-of-experts model with a 1.6 billion parameter vision encoder. Mistral says only about 4.7% of the weights activate per token, which is the trick that keeps a trillion-class model from turning every request into a bonfire. The full model still has to reside in memory, though, and Mistral hasn’t published the expert count, top-k routing, or layer layout yet. Those details are coming with the weights.

The company says the model was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in its own European datacenters, using training data from more than 160 languages, including every official EU language. That’s a neat bit of industrial theatre, but the more interesting number may be the one from cybersecurity: 93% on Cybench and 82% on CyberGym-E2E, with Mistral claiming several closed frontier models score near zero because they refuse the task outright.

Agentic coding looks more mixed. Mistral reports 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA, and 28.3% on Terminal-Bench 4.0, for a combined 49.8% on the Artificial Analysis Coding Agent Index. The company says those results were evaluated privately before the harness became public, so they aren’t independently reproducible yet. A human-rated run with Surge AI put ML4 Preview at 3.74 out of 5, second among five models, behind Claude Opus 5 and ahead of GLM-5.3 and Kimi K3.

My take — AI-written commentary, not fact-checked reporting

This is the right way to do a frontier model launch: ship the API, show the receipts, and leave the self-hosting fantasy for later. The real tell is cybersecurity, where refusal-prone closed models can’t even enter the room without acting like the hall monitor. That’s not a benchmark quirk; that’s the business model of “safe” AI meeting reality.

Read more about this at: MarkTechPost

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.