TLDRocket
Sign in

Introducing Qwen-VL

Qwen Covered by 2 sources

Alibaba released upgraded versions of its Qwen-VL multimodal model, introducing Qwen-VL-Plus and Qwen-VL-Max with improved image reasoning and text recognition capabilities. The new versions support image resolutions above one million pixels and various aspect ratios, addressing generalization limitations in multimodal AI systems. These upgrades enable better performance on tasks requiring detailed image analysis and text extraction from images.

Why it matters

Along with the rapid development of our large language model Qwen, we leveraged Qwen’s capabilities and unified multimodal pretraining to address the limitations of multimodal models in generalization, and we opensourced multimodal model Qwen-VL in Sep. 2023. Recently, the Qwen-VL series has undergone a significant upgrade with the launch of two enhanced versions, Qwen-VL-Plus and Qwen-VL-Max. The key technical advancements in these versions include: Substantially boost in image-related reasoning capabilities; Considerable enhancement in recognizing, extracting, and analyzing details within images and texts contained therein; Support for high-definition images with resolutions above one million pixels and images of various aspect ratios.

Also covered by

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.