Qwen VLo: From "Understanding" the World to "Depicting" It
Qwen
Alibaba's Qwen released Qwen VLo, a unified multimodal model that both understands and generates images, supporting image editing through natural language instructions like style transfer and object modification. The model uses progressive left-to-right, top-to-bottom generation and supports dynamic image resolutions with aspect ratios as extreme as 4:1, currently available as a preview in Qwen Chat. The addition of generative capabilities enables more flexible creative workflows and allows the model to verify its own understanding through intermediate outputs like segmentation and detection maps.
Why it matters
QWEN CHAT DISCORD Introduction The evolution of multimodal large models is continually pushing the boundaries of what we believe technology can achieve. From the initial QwenVL to the latest Qwen2.5 VL, we have made progress in enhancing the model’s ability to understand image content. Today, we are excited to introduce a new model, Qwen VLo, a unified multimodal understanding and generation model. This newly upgraded model not only “understands” the world but also generates high-quality recreations based on that understanding, truly bridging the gap between perception and creation.