TLDRocket
Sign in

SpaceXAI Releases Grok Voice Transcribe 2.0: A Speech-to-Text API Claiming 2x Accuracy Over 1.0 at $0.10 per Hour

MarkTechPost Michal Sutter

SpaceXAI launched Grok Voice Transcribe 2.0, a speech-to-text API. It claims 2x the accuracy of 1.0 for the same price, but only as hosted API.

Based on reporting by MarkTechPost, Michal Sutter — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

SpaceXAI has shipped Grok Voice Transcribe 2.0, a new speech-to-text model that lives behind its API and nothing else. The company says it is twice as accurate as Grok Voice Transcribe 1.0 while keeping pricing unchanged, and it is already live under the model ID grok-voice-transcribe-2.0.

The pitch is pretty specific: rough audio. Think noisy phone lines, overlapping speakers, accents, and people rattling off emails or phone numbers. SpaceXAI says the model runs in both batch and streaming modes, and that it can also detect language automatically, then keep going if the speaker switches languages mid-recording.

Under the hood, the model comes from the audio foundation model behind Grok Voice. SpaceXAI says that system already handles tens of thousands of customer-support calls a day, transcribes millions of hours of video narration, and powers the Grok assistant in Tesla vehicles. The training data, according to the company, was live, noisy, multilingual audio gathered from varied environments, followed by post-training refinement.

The benchmark claims are the headline here. SpaceXAI says the model ranks first among 32 streaming models on the public Artificial Analysis leaderboard. It also reports internal word-error-rate tests across four production-style sets: telephony, conversational speech, credentials, and short phrases in 19 languages. The biggest jump was on short phrases, where WER fell from 20.6% to 6.8%.

For developers, the API bundles a lot into one place: word timestamps, speaker diarization, multichannel transcription, key-term biasing, text formatting, filler-word removal, and turn detection. Batch jobs accept files up to 500 MB across 12 audio formats. Pricing stays at $0.10 per audio hour for batch and $0.20 for streaming, with diarization and timestamps included. Atlassian Loom is already using it to transcribe every video, which is a decent proof point even if the vendor claims still deserve a skeptical eyebrow.

My take — AI-written commentary, not fact-checked reporting

API-only speech tools are the right bet here. Most teams don’t want to babysit weights; they want fewer bad transcripts and a bill they can explain. The open-weight crowd can keep their purity tests — the market keeps rewarding whatever gets the words right.

Read more about this at: MarkTechPost

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.