TLDRocket
Sign in

Qwen2-Audio: Chat with Your Voice!

Qwen

Alibaba released Qwen2-Audio, an updated version of its multimodal AI model that accepts audio and text inputs and generates text outputs. The model builds on Qwen-VL and previous versions of Qwen-Audio to extend the company's large language model capabilities across multiple modalities. This advancement enables users to interact with the AI system using voice input alongside text, expanding the range of input methods for the model.

Why it matters

DEMO PAPER GITHUB HUGGING FACE MODELSCOPE DISCORD To achieve the objective of building an AGI system, the model should be capable of understanding information from different modalities. Thanks to the rapid development of large language models, LLMs are now capable of understanding language and reasoning. Previously we have taken a step forward to extend our LLM, i.e., Qwen, to more modalities, including vision and audio, and built Qwen-VL and Qwen-Audio. Today, we release Qwen2-Audio, the next version of Qwen-Audio, which is capable of accepting audio and text inputs and generating text outputs.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.