Microsoft AI Releases MAI-Transcribe-2-Streaming: #1 Real-Time Speech-to-Text Model on Artificial Analysis
MarkTechPost Michal Sutter ● Covered by 5 sources
Microsoft launched its first streaming speech-to-text model, MAI-Transcribe-2-Streaming. It’s topping a key accuracy test while giving apps answers before the speaker is done.
Based on reporting by MarkTechPost, Michal Sutter — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Microsoft AI has rolled out MAI-Transcribe-2-Streaming, its first streaming speech-to-text model, alongside two text-to-speech models: MAI-Voice-2.1 and MAI-Voice-2.1-Flash. The timing matters. This is the company’s move into real-time transcription, where the difference between hearing a sentence and seeing it on screen can decide whether an assistant feels sharp or sluggish.
The model is the streaming counterpart to MAI-Transcribe-2, which arrived in September. It supports 60 languages and keeps detecting the language automatically as audio keeps coming in. Microsoft says the model can start emitting partial transcripts just over 100ms after receiving audio, then keep revising them until it settles on a final version. That gives voice agents something important: they do not have to wait for the speaker to stop before they start acting.
Artificial Analysis puts MAI-Transcribe-2-Streaming at the top of its AA-WER Streaming ranking, ahead of 37 other models. In that benchmark, the model posted 2.5% word error rate for both its final transcript and its first partial transcript. It did that at 0.13 seconds after the end of speech for the final, and 0.12 seconds for the first partial. The next closest runners-up were Grok Voice Transcribe 2.0 at 2.7% and 0.49 seconds, and Muse Voice Transcribe at 3.1% and 0.16 seconds.
Speed is not the whole story. Cartesia Ink-2 returns final results faster at 0.07 seconds, but with a worse 4.0% WER. Microsoft’s own pitch is that its first partial is already as accurate as the final transcript, which is exactly the kind of detail that matters if a system is going to trigger tools before a person finishes talking.
Pricing is aggressive only in the abstract. Microsoft is charging $0.54 per hour of audio as an introductory rate through the end of 2026, which Artificial Analysis normalizes to $9.00 per 1,000 minutes. The batch MAI-Transcribe-2 model is much cheaper at $0.10 per hour. On the streaming side, Microsoft is pricier than xAI and Meta, and roughly in line with Google’s estimated rate. Developers can wire it up through a Realtime API or the Azure Speech SDK, and it is also available in the MAI Playground, through Vercel, and in Azure Voice Live. LiveKit support is listed as coming soon.
My take — AI-written commentary, not fact-checked reporting
This is the part of AI that actually counts: not bigger demos, but less waiting. A lot of vendors still sell “real time” like a slogan; Microsoft is trying to make it a measurable product, which is a healthier way to compete. The awkward bit is that good latency is still being sold at a premium, because of course it is.
Read more about this at: MarkTechPost
Related stories
Google AI Releases Gemini 3.5 Transcribe: A Speech-to-Text Model Reporting 2.6% Average WER Across 85+ Languages
MarkTechPost · 1 month ago ·
43
Meta Superintelligence Labs Releases Muse Voice Transcribe: One Real-Time Model for Streaming ASR, Diarization, and Endpointing
MarkTechPost · 1 month ago ·
5