Alibaba’s Tongyi Lab Releases Qwen-Audio-3.0-TTS, a Hosted Text-to-Speech Model in Flash and Plus Tiers Across 16 Languages
MarkTechPost 1 month ago 50
Alibaba's Tongyi Lab released Qwen-Audio-3.0-TTS, a hosted text-to-speech model available in two tiers (Flash for real-time interaction and Plus for high quality) supporting 16 languages. The Plus variant ranks first on the Artificial Analysis leaderboard with an Elo rating near 1,236 and costs $27.59 per million characters. The model includes 86 fine-grained inline tags for controlling non-verbal details like laughter and breathing, but is available only as a hosted API rather than downloadable weights.