Un Ministral, des Ministraux
Mistral AI
Mistral just dropped two tiny new models built for phones, gadgets, and offline devices. One of them, at just 3B params, already beats last year's flagship 7B model on benchmarks.
Based on reporting by Mistral AI — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
A year to the day after Mistral 7B first shook up the open-model scene, the French lab is back with something smaller and, in its own telling, sharper. Ministral 3B and Ministral 8B are the new entries, aimed squarely at edge computing: phones, robots, local devices, anything that needs to think without phoning home to a data center.
The pitch is privacy and latency. Mistral says customers have been pushing for on-device translation, assistants that work without internet, local analytics, and robotics that don't lean on cloud inference. Both Ministraux support context windows up to 128k tokens, though vLLM currently caps that at 32k, and the 8B variant adds an interleaved sliding-window attention setup meant to keep memory use down while speeding up inference.
What's notable is the claim that Ministral 3B, despite being less than half the size of last year's Mistral 7B, now beats it on most benchmarks. Mistral ran its own internal evaluation framework against rivals like Gemma 2, Llama 3.2, and Llama 3.1, and says both new models come out ahead in their weight classes across knowledge, reasoning, and function-calling tasks. That last one matters for a specific reason: Mistral is positioning these small models as the workhorses inside larger agentic systems, handling input parsing, routing, and API calls while a bigger model like Mistral Large handles the heavy thinking.
Pricing is aggressive. Ministral 8B runs $0.10 per million tokens, Ministral 3B just $0.04, both live now on la Plateforme. Weights for the 8B Instruct version are open for research use, and Mistral says it'll help enterprises quantize the models for self-hosted deployment under its commercial license. Cloud partners are expected to carry both models soon.
My take — AI-written commentary, not fact-checked reporting
Small models beating last year's flagship in twelve months is the real story of 2024, not another trillion-parameter monster nobody can run locally. I'd rather see labs compete on efficiency than on who can burn the most GPU hours, and Mistral keeping an 8B model's weights open for research is exactly the kind of move that keeps European AI relevant instead of just watching OpenAI and Google duke it out.
Read more about this at: Mistral AI