TLDRocket
Sign in

WAXAL: A large-scale open resource for African language speech technology

Google Research

Google just open-sourced a massive speech dataset covering 27 African languages spoken by 100M+ people. Voice tech has ignored most of the world's languages for years — this is a real dent in that gap.

Based on reporting by Google Research — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

For all the noise about AI understanding human speech, most of that understanding stops at a handful of languages. English, Mandarin, Spanish — the usual suspects. Meanwhile Sub-Saharan Africa, home to more than 2,000 distinct languages, has largely been locked out of voice assistants, transcription tools, and the rest of the speech-tech stack. Google Research has been chipping away at that problem since 2021, and the result is WAXAL, a dataset now public under a CC-BY-4.0 license covering 27 languages across more than 26 countries.

The numbers are substantial: roughly 1,846 hours of transcribed spontaneous speech for ASR, plus over 565 hours of studio-grade audio for TTS. What's notable isn't just the volume but the method. Instead of having people read scripts aloud — which tends to produce stiff, unnatural speech — the ASR collection had participants describe images across more than 50 topics in their own languages. That approach captured tonal variation and code-switching, the messy stuff that actually happens when people talk, rather than the clean but artificial speech you get from script-reading. For the TTS side, local pairs drafted scripts of 10,000 to 20,000 words and took turns recording each other, with some teams building their own makeshift studio booths to get broadcast-quality sound.

The more interesting part of this story is who did the work. Google didn't just fund a data-collection exercise and slap its name on it. Makerere University handled nine languages, the University of Ghana covered eight, and groups like Digital Umuganda, Addis Ababa University, Media Trust, Loud n Clear, and AIMS Senegal split the rest. Crucially, partners retained ownership of their data even as it went into the open pool — a structure that matters if you want this kind of effort to be repeatable rather than a one-off extraction.

The dataset has already spun off real research: a cookbook for collecting speech from people with conditions like cerebral palsy and stammering, which produced the first open dataset of Akan speakers with speech impairments; a 5,000-hour corpus spanning five Ghanaian languages; and a benchmarking study testing Whisper, XLS-R, MMS, and W2v-BERT across 13 African languages, which found that throwing more training data at a model helps unevenly depending on how tonal or morphologically complex the language is. A separate literature review catalogued 74 existing datasets across 111 African languages, essentially mapping how much ground is still uncovered.

WAXAL isn't a finished product — it's explicitly framed as a growing collection, with Google committing to keep adding languages. Given how badly skewed voice tech has been toward high-resource languages, that's the right instinct. Whether it actually gets picked up by developers building real products, versus just cited in papers, is the part nobody can promise yet.

My take — AI-written commentary, not fact-checked reporting

This is the kind of open-data project that actually deserves the label, since Google gave up ownership claims and let African institutions keep control of what they built — that's rarer than it should be from a company this size. I'd rather see ten of these than another closed frontier-model launch, because the real bottleneck in AI accessibility has never been model architecture, it's been who gets represented in the training data at all. Watch whether anyone actually ships a product on top of it within a year; datasets that just sit in papers don't move the needle.

Read more about this at: Google Research

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.