TLDRocket
Sign in

Meta just beat OpenAI and Google at real-time transcription

The New Stack Frederic Lardinois Covered by 7 sources

Meta launched a real-time speech model that beats OpenAI and Google on some benchmarks. It can handle 20+ speakers, 70+ languages, and long calls, but Meta won’t open-source it.

Based on reporting by The New Stack, Frederic Lardinois — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Meta’s Superintelligence Labs rolled out Muse Voice Transcribe on Tuesday, and the pitch is simple: this is the company’s first real-time audio perception model, and on some tests it is the best one around for live speech recognition. Meta is putting it into the Meta Model API, Meta AI for Mac, and Muse Code right away. It is not, however, sharing open weights, a spokesperson told The New Stack.

The benchmark results are the headline-grabber. On Artificial Analysis’s AA-WER Streaming speech-to-text benchmark in English, Muse Voice Transcribe posts a 3.1% word error rate. That edges out Cartesia Ink-2 at 3.4%, ElevenLabs’ Scribe v2 Realtime at 3.6%, GPT Live Transcribe at 3.9%, and Gemini 3.5 Transcribe Live at 4%. On speaker recognition, Meta says the model also leads the pack in real-time use cases with a 17.5% error rate across several standard benchmarks.

The model’s feature list is broad. Meta says it can distinguish more than 20 speakers, has been trained on more than 70 languages, and has 25 languages that were “extensively verified.” It can also handle multilingual speakers switching languages mid-conversation, and it supports conversations that run for over an hour.

Under the hood, Muse Voice Transcribe is an autoregressive multimodal model from the Muse Spark family. Audio arrives in 80-millisecond chunks, each compressed into a single soft token. For every chunk, the model either emits text or a special “next audio” placeholder, which gets swapped for the next slice of sound. When the audio ends, an “empty audio” token tells it to flush the text it is still holding.

That setup is what Meta calls “adaptive delay.” The model decides how much context it wants before committing to a word, so tricky bits can wait while easy ones get written almost immediately. The company says that tradeoff is learned during reinforcement learning, with word error rate and delay rewards multiplied together. And because real-time transcription is turning into a crowded race this summer, with OpenAI, Google, xAI, and Alibaba all shipping streaming models within weeks of each other, Meta clearly wants an edge it can keep improving for its glasses and Mac app.

My take — AI-written commentary, not fact-checked reporting

This is the usual Meta move: ship something genuinely strong, then keep the best version behind the wall. The company talks up openness when it suits the story, and locks the door when the model matters. For anyone watching the AI race, the real tell is not the benchmark win; it’s that Meta thinks this is good enough to protect.

Read more about this at: The New Stack

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.