Qwen2-Audio: Chat with Your Voice!
Qwen
Alibaba released Qwen2-Audio, an updated version of its multimodal AI model that accepts audio and text inputs and generates text outputs. The model builds on Qwen-VL and previous versions of Qwen-Audio to extend the company's large language model capabilities across multiple modalities. This advancement enables users to interact with the AI system using voice input alongside text, expanding the range of input methods for the model.
Why it matters
DEMO PAPER GITHUB HUGGING FACE MODELSCOPE DISCORD To achieve the objective of building an AGI system, the model should be capable of understanding information from different modalities. Thanks to the rapid development of large language models, LLMs are now capable of understanding language and reasoning. Previously we have taken a step forward to extend our LLM, i.e., Qwen, to more modalities, including vision and audio, and built Qwen-VL and Qwen-Audio. Today, we release Qwen2-Audio, the next version of Qwen-Audio, which is capable of accepting audio and text inputs and generating text outputs.