QVQ-Max: Think with Evidence
Qwen
Alibaba's Qwen team released QVQ-Max, a visual reasoning model that analyzes images and videos to solve problems across mathematics, coding, and creative tasks. The model shows continuous accuracy improvements on the MathVision benchmark as its reasoning process length increases, demonstrating enhanced problem-solving capabilities. The release enables broader applications in workplace productivity, education, and daily life assistance through improved visual understanding and analytical reasoning.
Why it matters
QWEN CHAT GITHUB HUGGING FACE MODELSCOPE DISCORD Introduction Last December, we launched QVQ-72B-Preview as an exploratory model, but it had many issues. Today, we are officially releasing the first version of QVQ-Max, our visual reasoning model. This model can not only “understand” the content in images and videos but also analyze and reason with this information to provide solutions. From math problems to everyday questions, from programming code to artistic creation, QVQ-Max has demonstrated impressive capabilities.