TLDRocket
Sign in

Microsoft targets ultra-realistic voice agents with its first streaming transcription model

SiliconANGLE Mike Wheatley ● Covered by 4 sources

Microsoft built a streaming speech model that updates transcripts as people talk. It’s meant to make voice agents feel less robotic, and it comes with three MAI models in one go.

Based on reporting by SiliconANGLE, Mike Wheatley — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Microsoft has rolled out a new batch of MAI models, and the headline act is its first streaming transcription model. The company is aiming squarely at voice agents that can listen, react and keep the conversation moving without awkward pauses. That’s the pitch, anyway: less lag, less waiting, more human-sounding back-and-forth.

The new transcription model is called MAI-Transcribe-2-Streaming. It takes speech through a WebSocket, keeps updating the transcript while someone is still talking, and then locks it in once the speaker stops. Microsoft says that makes it useful for live captions and for agents that can start working on a request before the user has fully finished speaking. The model is on Microsoft’s Vercel AI Gateway and costs 54 cents per audio hour.

Microsoft says it supports more than 60 languages and can identify the language automatically. It also produces its first transcript hypotheses in 320 milliseconds on average, though the company is careful not to promise that speed everywhere, since network conditions and the downstream AI system both matter. That caveat is doing a lot of work. Real-time AI always sounds cleaner in a demo than it does on a flaky connection.

The company is not stopping at transcription. It also introduced MAI-Voice-2.1 and MAI-Voice-2.1-Flash, two text-to-speech models for the output side of the conversation. The first is aimed at more expressive, higher-fidelity speech. The Flash version trades some of that polish for lower cost and faster responses. According to Microsoft’s Vercel listings, MAI-Voice-2.1 costs $22 per million characters and the Flash model costs $15 per million characters. Both support 23 languages.

The bigger strategic move is hard to miss. Microsoft is increasingly building around its own MAI family instead of leaning so heavily on OpenAI and Anthropic, even though it is a major investor in both. Mustafa Suleyman has already said the company wants to reduce and eventually eliminate what it pays Anthropic. Put bluntly, Microsoft wants its own voice stack, its own reasoning model, and its own bill.

My take — AI-written commentary, not fact-checked reporting

This is the sensible way to build voice AI: split the job into pieces and stop pretending one giant model should do everything. Microsoft’s real message is about control and cost, not just clever speech demos. The age of renting intelligence forever is starting to look expensive, which is an inconvenient fact for the hype machine.

Read more about this at: SiliconANGLE

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.