TLDRocket
Sign in

Microsoft AI Releases MAI-Transcribe-2-Streaming: #1 Real-Time Speech-to-Text Model on Artificial Analysis

MarkTechPost Michal Sutter ● Covered by 5 sources

Microsoft launched its first streaming speech-to-text model, MAI-Transcribe-2-Streaming. It’s topping a key accuracy test while giving apps answers before the speaker is done.

Based on reporting by MarkTechPost, Michal Sutter — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Microsoft AI has rolled out MAI-Transcribe-2-Streaming, its first streaming speech-to-text model, alongside two text-to-speech models: MAI-Voice-2.1 and MAI-Voice-2.1-Flash. The timing matters. This is the company’s move into real-time transcription, where the difference between hearing a sentence and seeing it on screen can decide whether an assistant feels sharp or sluggish.

The model is the streaming counterpart to MAI-Transcribe-2, which arrived in September. It supports 60 languages and keeps detecting the language automatically as audio keeps coming in. Microsoft says the model can start emitting partial transcripts just over 100ms after receiving audio, then keep revising them until it settles on a final version. That gives voice agents something important: they do not have to wait for the speaker to stop before they start acting.

Artificial Analysis puts MAI-Transcribe-2-Streaming at the top of its AA-WER Streaming ranking, ahead of 37 other models. In that benchmark, the model posted 2.5% word error rate for both its final transcript and its first partial transcript. It did that at 0.13 seconds after the end of speech for the final, and 0.12 seconds for the first partial. The next closest runners-up were Grok Voice Transcribe 2.0 at 2.7% and 0.49 seconds, and Muse Voice Transcribe at 3.1% and 0.16 seconds.

Speed is not the whole story. Cartesia Ink-2 returns final results faster at 0.07 seconds, but with a worse 4.0% WER. Microsoft’s own pitch is that its first partial is already as accurate as the final transcript, which is exactly the kind of detail that matters if a system is going to trigger tools before a person finishes talking.

Pricing is aggressive only in the abstract. Microsoft is charging $0.54 per hour of audio as an introductory rate through the end of 2026, which Artificial Analysis normalizes to $9.00 per 1,000 minutes. The batch MAI-Transcribe-2 model is much cheaper at $0.10 per hour. On the streaming side, Microsoft is pricier than xAI and Meta, and roughly in line with Google’s estimated rate. Developers can wire it up through a Realtime API or the Azure Speech SDK, and it is also available in the MAI Playground, through Vercel, and in Azure Voice Live. LiveKit support is listed as coming soon.

My take — AI-written commentary, not fact-checked reporting

This is the part of AI that actually counts: not bigger demos, but less waiting. A lot of vendors still sell “real time” like a slogan; Microsoft is trying to make it a measurable product, which is a healthier way to compete. The awkward bit is that good latency is still being sold at a premium, because of course it is.

Read more about this at: MarkTechPost

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.