QVQ-Max: Think with Evidence
GitHub Pages
Qwen dropped QVQ-Max, a visual reasoning model that looks at images and video, then thinks through problems step by step. It's an early step toward AI that reasons about what it sees, not just describes it.
Based on reporting by GitHub Pages — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Alibaba's Qwen team has quietly moved past its December preview and shipped what it calls the first real version of QVQ-Max, a model built to reason about images and video rather than just caption them. The pitch is straightforward: most AI still leans on text, but a huge chunk of the world's information — blueprints, charts, wardrobe photos, geometry diagrams — never gets translated into words cleanly. QVQ-Max is meant to look at that visual mess and actually think about it.
The team frames the model's skillset in three layers. First, it spots detail — objects, labels, small things a person might skim past in a photo. Second, and more interesting, it reasons from what it sees, combining visual input with background knowledge to solve a geometry problem or guess what happens next in a video clip. Third, it gets creative, turning a rough sketch into a finished illustration or spinning a photo into a script, a critique, even a fortune-telling bit.
Qwen backs this up with a benchmark result worth noting: on MathVision, a test suite of gnarly multimodal math problems, accuracy climbed steadily as they let the model's internal
My take — AI-written commentary, not fact-checked reporting
properly. Still, this is explicitly a version one. Qwen says the roadmap includes grounding techniques to cut down on hallucinated observations, an
Read more about this at: GitHub Pages