TLDRocket
Sign in

Voxtral

Mistral AI

Mistral just dropped Voxtral, two open speech-AI models that transcribe and understand audio. They beat Whisper and undercut ElevenLabs on price, and you can run them yourself.

Based on reporting by Mistral AI — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Mistral has a habit of releasing things under Apache 2.0 right when everyone assumes the good stuff stays locked behind an API, and Voxtral is another example of that pattern. The French lab shipped two speech models this week — a 24B version aimed at production workloads and a 3B version small enough to run on a laptop or edge device — and both are free to download from Hugging Face, no strings attached.

What makes Voxtral interesting isn't just that it transcribes speech well, though it does: Mistral claims it beats Whisper large-v3 across the board and edges out GPT-4o mini Transcribe and Gemini 2.5 Flash on several benchmarks, including English short-form audio and Mozilla Common Voice. It's the fact that these models skip the usual pipeline entirely. Instead of bolting a language model onto a separate transcription engine, Voxtral handles the whole chain itself — you can hand it a 40-minute audio file and ask it to summarize the meeting, answer a question about what someone said at minute 12, or trigger a function call based on spoken intent, all without stitching together two different systems.

The pricing is the other half of the pitch. Mistral says Voxtral Mini Transcribe beats Whisper for less than half the cost, and Voxtral Small matches ElevenLabs Scribe at a similar discount, with API access starting at a tenth of a cent per minute. For anyone building voice features into a product — call center tooling, meeting assistants, multilingual customer support — that's the kind of price gap that changes which projects are worth greenlighting.

There's also a multilingual angle worth flagging: Voxtral auto-detects language and performs well across Spanish, French, Hindi, German, Dutch, Portuguese and Italian, which matters more than most benchmark charts suggest if you're trying to ship a product outside the English-speaking world. Mistral built it on top of Mistral Small 3.1, so it keeps the text-model chops too, meaning it can double as a general-purpose language model when you're not feeding it audio.

None of this is finished, by Mistral's own admission. Speaker diarization, emotion tagging, word-level timestamps and non-speech audio detection are all listed as coming later, and the company is openly recruiting an audio research team to get there. But shipping a capable open model now, with a clear roadmap and enterprise fine-tuning options already available, is a stronger opening move than most labs manage.

My take — AI-written commentary, not fact-checked reporting

I like that Mistral keeps forcing the closed-API crowd to justify their prices, and undercutting ElevenLabs by half while beating Whisper on accuracy is a genuinely good outcome for anyone who isn't a giant lab with unlimited compute. The real test isn't the benchmark chart, though — it's whether Voxtral's semantic layer holds up on messy real-world audio, accents, and overlapping speakers, where open models have historically fallen apart the moment the demo ends.

Read more about this at: Mistral AI

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.