Open TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice Cloning
Hugging Face
Open TTS Leaderboard ranks voice models with objective tests, not just votes. It’s meant to keep up with 8K+ TTS models and give open-source voices a fairer shot.
Based on reporting by Hugging Face — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Open-source text-to-speech has been moving fast. On Hugging Face Hub, there are now more than 8K TTS models, but the way people evaluate them has not kept up. The result is a mess of fragmented scoring, with arena-style leaderboards doing their best to catch a moving target.
Those arenas still matter. They let users listen to two outputs, pick a favorite, and turn enough votes into an Elo ranking. But they’re slow, and they skew toward models that are easy to plug in. On Artificial Analysis, only 16 of 92 models are open-weights as of Sep. 30, 2026, with a similar pattern on Voice Arena. Hosted API models are simpler to add; open models need to be run and served by the arena operator.
Hugging Face’s Open TTS Leaderboard takes a different route. It leans on objective metrics: intelligibility through WER and CER using Qwen3 ASR, speed through RTFx and time-to-first-audio, and voice similarity through WavLM speaker embeddings. The team says this cuts evaluation from weeks of vote collection to a couple of hours. That is a real shift, even if it doesn’t pretend to settle everything.
And it doesn’t try to. The leaderboard is explicit that ASR-based WER is only a proxy for intelligibility, while speaker similarity only approximates identity preservation. Naturalness, expressiveness, and plain human preference still need people in the loop. That’s why the site also has a Listen tab, where users can compare outputs directly and vote with an HF account.
The ranking views are built to surface different tradeoffs. The default view focuses on English WER across Seed TTS Eval and CV3 Eval. A multilingual mode changes the picture, and a voice-cloning toggle adds speaker similarity plus Pareto plots for SIM, batched inference, and model size. There is also a Streaming tab, which measures time to first audio on the same 50 English prompts from CV3-Eval, on H200 GPU and, for some models, CPU. The point is clear: open TTS needs more than applause. It needs measurement.
My take — AI-written commentary, not fact-checked reporting
This is the right kind of boring. TTS has spent too long pretending vote counts alone can carry the whole job, which is a nice story until open models get squeezed out by whatever is easiest to host. Objective metrics won’t replace human ears, but they do stop the leaderboard from becoming a beauty contest with a server bill.
Read more about this at: Hugging Face
Related stories
Best Voice Cloning APIs in 2026: Speaker Similarity, Consent Checks, and Price per 1M Characters
MarkTechPost · 1 week ago ·
9
Introducing the FFASR Leaderboard: Benchmarking ASR in the Real World
Hugging Face · 3 months ago ·
32