TLDRocket
Sign in

Open TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice Cloning

Hugging Face

Open TTS Leaderboard ranks voice models with objective tests, not just votes. It’s meant to keep up with 8K+ TTS models and give open-source voices a fairer shot.

Based on reporting by Hugging Face — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Open-source text-to-speech has been moving fast. On Hugging Face Hub, there are now more than 8K TTS models, but the way people evaluate them has not kept up. The result is a mess of fragmented scoring, with arena-style leaderboards doing their best to catch a moving target.

Those arenas still matter. They let users listen to two outputs, pick a favorite, and turn enough votes into an Elo ranking. But they’re slow, and they skew toward models that are easy to plug in. On Artificial Analysis, only 16 of 92 models are open-weights as of Sep. 30, 2026, with a similar pattern on Voice Arena. Hosted API models are simpler to add; open models need to be run and served by the arena operator.

Hugging Face’s Open TTS Leaderboard takes a different route. It leans on objective metrics: intelligibility through WER and CER using Qwen3 ASR, speed through RTFx and time-to-first-audio, and voice similarity through WavLM speaker embeddings. The team says this cuts evaluation from weeks of vote collection to a couple of hours. That is a real shift, even if it doesn’t pretend to settle everything.

And it doesn’t try to. The leaderboard is explicit that ASR-based WER is only a proxy for intelligibility, while speaker similarity only approximates identity preservation. Naturalness, expressiveness, and plain human preference still need people in the loop. That’s why the site also has a Listen tab, where users can compare outputs directly and vote with an HF account.

The ranking views are built to surface different tradeoffs. The default view focuses on English WER across Seed TTS Eval and CV3 Eval. A multilingual mode changes the picture, and a voice-cloning toggle adds speaker similarity plus Pareto plots for SIM, batched inference, and model size. There is also a Streaming tab, which measures time to first audio on the same 50 English prompts from CV3-Eval, on H200 GPU and, for some models, CPU. The point is clear: open TTS needs more than applause. It needs measurement.

My take — AI-written commentary, not fact-checked reporting

This is the right kind of boring. TTS has spent too long pretending vote counts alone can carry the whole job, which is a nice story until open models get squeezed out by whatever is easiest to host. Objective metrics won’t replace human ears, but they do stop the leaderboard from becoming a beauty contest with a server bill.

Read more about this at: Hugging Face

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.