Alibaba’s Tongyi Lab Releases Qwen-Audio-3.0-TTS, a Hosted Text-to-Speech Model in Flash and Plus Tiers Across 16 Languages
MarkTechPost Asif Razzaq
Alibaba's Tongyi Lab released Qwen-Audio-3.0-TTS, a hosted text-to-speech model available in two tiers (Flash for real-time interaction and Plus for high quality) supporting 16 languages. The Plus variant ranks first on the Artificial Analysis leaderboard with an Elo rating near 1,236 and costs $27.59 per million characters. The model includes 86 fine-grained inline tags for controlling non-verbal details like laughter and breathing, but is available only as a hosted API rather than downloadable weights.
Why it matters
Alibaba’s Tongyi Lab has released Qwen-Audio-3.0-TTS, a production-oriented text-to-speech (TTS) system. The model ships in two variants from the same lineage. Flash targets real-time interaction. Plus targets high-quality generation. Both are delivered as hosted models through Alibaba Cloud Model Studio, not as downloadable weights. The release focuses on four things developers hit in production: broader […] The post Alibaba’s Tongyi Lab Releases Qwen-Audio-3.0-TTS, a Hosted Text-to-Speech Model in Flash and Plus Tiers Across 16 Languages appeared first on MarkTechPost.