Microsoft targets ultra-realistic voice agents with its first streaming transcription model
SiliconANGLE Mike Wheatley ● Covered by 4 sources
Microsoft built a streaming speech model that updates transcripts as people talk. It’s meant to make voice agents feel less robotic, and it comes with three MAI models in one go.
Based on reporting by SiliconANGLE, Mike Wheatley — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Microsoft has rolled out a new batch of MAI models, and the headline act is its first streaming transcription model. The company is aiming squarely at voice agents that can listen, react and keep the conversation moving without awkward pauses. That’s the pitch, anyway: less lag, less waiting, more human-sounding back-and-forth.
The new transcription model is called MAI-Transcribe-2-Streaming. It takes speech through a WebSocket, keeps updating the transcript while someone is still talking, and then locks it in once the speaker stops. Microsoft says that makes it useful for live captions and for agents that can start working on a request before the user has fully finished speaking. The model is on Microsoft’s Vercel AI Gateway and costs 54 cents per audio hour.
Microsoft says it supports more than 60 languages and can identify the language automatically. It also produces its first transcript hypotheses in 320 milliseconds on average, though the company is careful not to promise that speed everywhere, since network conditions and the downstream AI system both matter. That caveat is doing a lot of work. Real-time AI always sounds cleaner in a demo than it does on a flaky connection.
The company is not stopping at transcription. It also introduced MAI-Voice-2.1 and MAI-Voice-2.1-Flash, two text-to-speech models for the output side of the conversation. The first is aimed at more expressive, higher-fidelity speech. The Flash version trades some of that polish for lower cost and faster responses. According to Microsoft’s Vercel listings, MAI-Voice-2.1 costs $22 per million characters and the Flash model costs $15 per million characters. Both support 23 languages.
The bigger strategic move is hard to miss. Microsoft is increasingly building around its own MAI family instead of leaning so heavily on OpenAI and Anthropic, even though it is a major investor in both. Mustafa Suleyman has already said the company wants to reduce and eventually eliminate what it pays Anthropic. Put bluntly, Microsoft wants its own voice stack, its own reasoning model, and its own bill.
My take — AI-written commentary, not fact-checked reporting
This is the sensible way to build voice AI: split the job into pieces and stop pretending one giant model should do everything. Microsoft’s real message is about control and cost, not just clever speech demos. The age of renting intelligence forever is starting to look expensive, which is an inconvenient fact for the hype machine.
Read more about this at: SiliconANGLE
Related stories
Meta Superintelligence Labs Releases Muse Voice Transcribe: One Real-Time Model for Streaming ASR, Diarization, and Endpointing
MarkTechPost · 4 weeks ago ·
5
Announcing the fastest inference for realtime voice AI agents
Together AI · 10 months ago ·
52
SpaceXAI Releases Grok Voice Transcribe 2.0: A Speech-to-Text API Claiming 2x Accuracy Over 1.0 at $0.10 per Hour
MarkTechPost · 1 week ago ·
8