Pixtral Large
Mistral AI ● Covered by 2 sources
Mistral just dropped Pixtral Large, a 124B open-weights model that reads charts, docs and images without dumbing down its text skills. It beats GPT-4o and Gemini on several vision benchmarks and it's open-weights, not locked behind an API.
Based on reporting by Mistral AI — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Mistral AI has a habit of releasing serious hardware right when you least expect it, and Pixtral Large is no exception. It's a 124-billion-parameter multimodal model built on top of Mistral Large 2, pairing a hefty 123B text decoder with a 1B vision encoder, and it's available with open weights under a research license. That combination matters because most vision-capable models trade away text quality to get there. Mistral says this one doesn't.
The numbers back up the pitch. Pixtral Large scores 69.4% on MathVista, a benchmark built around reasoning over charts and diagrams, putting it ahead of every model Mistral tested against. It also beats GPT-4o and Gemini 1.5 Pro on ChartQA and DocVQA, two benchmarks that test how well a model actually parses documents rather than just describing pictures. On the LMSys Vision Leaderboard, Pixtral Large sits nearly 50 ELO points clear of the next-best open-weights model, and it edges out GPT-4o's August build too. That's not a marginal win, that's a real gap.
The demo examples Mistral shared are the kind of practical stuff people actually want from a multimodal model: reading a Swiss café receipt and calculating an 18% tip in francs, spotting exactly where a training-loss chart spikes at the 10,000 and 20,000 step marks, and pulling company logos out of a slide to identify which businesses use Mistral's models. None of it is flashy AI-generated art. It's document literacy, and that's where enterprise money actually gets spent.
Mistral quietly upgraded Mistral Large itself alongside this launch, pushing out version 24.11 with better long-context handling, a refreshed system prompt, and more reliable function calling. That version will land on Google Cloud and Microsoft Azure within the week, and it's aimed squarely at RAG pipelines and agentic workflows rather than chatbot small talk. Pairing that release with Pixtral Large suggests Mistral wants a single coherent stack, text and vision, that enterprises can deploy without stitching together separate models from different vendors.
What's notable here isn't just the benchmark wins, it's the licensing. Mistral Research License for research use, a separate commercial license for production, and actual downloadable weights on Hugging Face. In a year where most frontier multimodal announcements come wrapped in API-only access, Mistral keeps betting that openness plus performance is a viable business, not just a research flex.
My take — AI-written commentary, not fact-checked reporting
I run a site obsessed with tracking who's actually shipping open weights versus who's just talking about openness, and Mistral keeps putting its money where its mouth is while Californian labs hide their best vision models behind paywalls. A European company beating GPT-4o on document understanding and letting you download the weights is exactly the kind of story that should worry OpenAI more than it currently seems to. My only gripe: 124B parameters isn't exactly something you're running on a laptop, so let's not pretend this is democratizing anything for the average developer just yet.
Read more about this at: Mistral AI