Qwen2.5 Omni: See, Hear, Talk, Write, Do It All!
Qwen
Alibaba released Qwen2.5-Omni, a multimodal AI model that processes text, images, audio, and video inputs while generating text and speech responses in real-time. The 7-billion-parameter model is available on Hugging Face, ModelScope, DashScope, and GitHub. The release expands Qwen's capabilities to handle diverse input and output formats simultaneously.
Why it matters
QWEN CHAT HUGGING FACE MODELSCOPE DASHSCOPE GITHUB PAPER DEMO DISCORD We release Qwen2.5-Omni, the new flagship end-to-end multimodal model in the Qwen series. Designed for comprehensive multimodal perception, it seamlessly processes diverse inputs including text, images, audio, and video, while delivering real-time streaming responses through both text generation and natural speech synthesis. To try the latest model, feel free to visit Qwen Chat and choose Qwen2.5-Omni-7B. The model is now openly available on Hugging Face, ModelScope, DashScope,and GitHub, with technical documentation available in our Paper.