Time to Speak Some Dialects, Qwen-TTS!
Qwen
Qwen released an updated text-to-speech model that supports generating speech in 3 Chinese dialects alongside standard Mandarin. The model was trained on over millions of hours of speech data and offers 7 bilingual voices across the supported dialects. The system automatically adjusts prosody, pacing, and emotional inflections based on input text for more natural output.
Why it matters
API DISCORD Introduction Here we introduce the latest update of Qwen-TTS (qwen-tts-latest or qwen-tts-2025-05-22) through Qwen API . Trained on a large-scale dataset encompassing over millions of hours of speech, Qwen-TTS achieves human-level naturalness and expressiveness. Notably, Qwen-TTS automatically adjusts prosody, pacing, and emotional inflections in response to the input text. Notably, Qwen-TTS supports the generation of 3 Chinese dialects, including Pekingese, Shanghainese, and Sichuanese. As of now, Qwen-TTS supports 7 Chinese-English bilingual voices, including Cherry, Ethan, Chelsie, Serena, Dylan (Pekingese), Jada (Shanghainese) and Sunny (Sichuanese).