TLDRocket
Sign in

Qwen2.5 VL! Qwen2.5 VL! Qwen2.5 VL!

GitHub Pages Covered by 3 sources

Alibaba released Qwen2.5-VL, a vision-language model available in 3B, 7B, and 72B parameter sizes that can recognize objects, analyze documents and charts, process videos over 1 hour long, and function as a visual agent. The 7B model outperforms GPT-4o-mini on multiple tasks while the 72B model achieves competitive performance with much larger models like Llama-3-405B-Instruct. The model improves efficiency through a redesigned visual encoder with window attention and adds capabilities like structured output generation for financial documents and second-level event localization in videos.

Why it matters

QWEN CHAT GITHUB HUGGING FACE MODELSCOPE DISCORD We release Qwen2.5-VL, the new flagship vision-language model of Qwen and also a significant leap from the previous Qwen2-VL. To try the latest model, feel free to visit Qwen Chat and choose Qwen2.5-VL-72B-Instruct. Also, we open both base and instruct models in 3 sizes, including 3B, 7B, and 72B, in both Hugging Face and ModelScope. The key features include: Understand things visually: Qwen2.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.