TLDRocket
Sign in

TTS Arena: Benchmarking Text-to-Speech Models in the Wild

Hugging Face

Hugging Face launched TTS Arena, a side-by-side voting site for text-to-speech models. It uses blind human votes, not shaky metrics like WER, to rank who actually sounds real.

Based on reporting by Hugging Face — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Text-to-speech has a measurement problem. Word error rate tells you if a model transcribed words correctly, not whether it sounds like a human or a Speak & Spell from 1985. Mean opinion score surveys exist too, but they're usually run on a handful of listeners in a lab, which makes them nearly useless for settling close calls between two decent models. Hugging Face's answer, launched this week, is to just ask a lot of people to listen and vote.

The tool is called TTS Arena, and it borrows its entire premise from LMSYS's Chatbot Arena, which has racked up more than 300,000 human rankings for language models. Same idea here, just for voices: you type in some text, two different models read it aloud, you listen to both, and you pick the one that sounds more natural. The catch, deliberately, is that you don't find out which model made which clip until after you vote — a small design choice meant to stop people from just clicking their favorite brand name.

Six models are in the arena at launch: ElevenLabs, the lone proprietary entry, alongside open-source options MetaVoice, OpenVoice, Pheme, WhisperSpeech, and XTTS. Hugging Face picked these specifically because they're considered the current best of what's publicly available, and because putting a closed commercial model like ElevenLabs next to open alternatives lets anyone see for themselves how far open-source TTS has actually come — or hasn't.

Votes feed into a leaderboard that ranks models with an Elo-style rating system, the same math chess uses to rank players. It starts empty and fills in as votes accumulate, updating continuously rather than being a fixed snapshot. That's a meaningfully different approach than a one-time benchmark paper: the ranking is alive, and it can shift as new models get added or as public taste in

My take — AI-written commentary, not fact-checked reporting

This is the right instinct — synthetic ears can't judge naturalness, so stop pretending WER numbers settle anything and just crowdsource the human ear instead. My only quibble is the model lineup skews toward whoever open-sourced fastest rather than whoever sounds best, and putting only one proprietary model in the ring feels like a token gesture rather than a real open-vs-closed test. Still, an Elo leaderboard beats another glossy MOS chart nobody can reproduce.

Read more about this at: Hugging Face

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.