TLDRocket
Sign in

Best Open Speech Recognition (ASR) Models in 2026: WER, Languages, Latency, and License Compared

MarkTechPost Asif Razzaq

Multiple open-source speech recognition models now compete at similar accuracy levels, with Cohere's Transcribe (5.42% WER), IBM's Granite Speech 4.1 (5.33%), and others within one percentage point of each other. The Open ASR Leaderboard rankings are unreliable because models are evaluated on different test sets—excluding easier benchmarks like TED-LIUM artificially inflates some scores. Model selection now depends on license type, language support, streaming capability, and cost per audio-hour rather than benchmark rank, making this a procurement decision rather than a research one.

Why it matters

Open speech recognition stopped being a Whisper monoculture in 2026. Cohere Transcribe, IBM Granite Speech 4.1, ARK-ASR and MOSS-Transcribe are now separated by less than one WER point on the Hugging Face Open ASR Leaderboard — which means rank no longer decides anything. This roundup compares 16 open-weight models on word error rate, language coverage, streaming latency and license, and shows why the published averages cannot be subtracted from one another. The post Best Open Speech Recognition (ASR) Models in 2026: WER, Languages, Latency, and License Compared appeared first on MarkTechPost.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.