TLDRocket
Sign in

Alibaba Qwen Releases Qwen-Audio-3.1-Realtime: A Full-Duplex Voice Model Trained to Think, Act, and Decide When to Speak

MarkTechPost Asif Razzaq ● Covered by 3 sources

Alibaba’s Qwen team launched Qwen-Audio-3.1-Realtime, a voice model that can think, act and decide when to speak. It’s API-only for now, and the pricing cuts are steep.

Based on reporting by MarkTechPost, Asif Razzaq — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Alibaba’s Qwen team has shipped Qwen-Audio-3.1, a five-model audio stack that covers speech recognition, text-to-speech, and realtime conversation. The headline piece is Qwen-Audio-3.1-Realtime, a full-duplex model aimed at voice agents that can call tools instead of just chatting prettily.

The service is already live as a managed API under qwen-audio-3.1-realtime-plus on QwenCloud, exposed over WebSocket. There are no open weights. Qwen is also pushing hard on price: the company says Realtime is about 85% cheaper than before, TTS about 70% cheaper, and ASR up to 95% cheaper.

The QwenCloud setup is fairly roomy. The model page lists 262K tokens of context, with 245K available for input and 16K for output. Default limits are 60 requests and 100K tokens per minute. Pricing is set at $6.4 per 1M audio input tokens, $0.8 per 1M text input tokens, and $24 per 1M output tokens for text and audio, with output text not charged.

This is not just a speech box. The stack includes function calling, web search, structured outputs, context cache, and fine-tuning. There is also Qwen-Audio-3.1-ASR-Flash-Filetrans for offline long-audio transcription, with hot words, speaker separation, punctuation, and multilingual plus Chinese dialect recognition.

Under the hood, Qwen describes a two-model design with shared Audio Encoder and LLM logic. One model decides whether to keep listening, speak, stop, or resume. Another writes the response as text. A voice renderer then turns that text into streaming speech, using conversation history, voice cues, and acoustic context.

The training setup is split into three layers: Think, Act, and Speak and Coordinate. In the Think stage, Qwen says it uses paired audio data at million-hour scale and on-policy distillation. In Act, the model learns inside executable environments with tool pools, JSON state, and business policies. In Speak and Coordinate, it learns when and how to answer, with some clear trade-offs: fewer unwanted replies to background talk, but a higher resume rate after interruptions and slower stop latency than GPT-Realtime-2.

My take — AI-written commentary, not fact-checked reporting

This is the right kind of ambition: less chatbot karaoke, more system that can actually decide when to shut up. The catch is the usual one — if the only thing you can self-host is disappointment, API-only “realtime” is just another rented brain with a nice brochure. Open weights would have made the story much more interesting.

Read more about this at: MarkTechPost

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.