TLDRocket
Sign in

Qwen

47 summarised stories about Qwen, each linking back to the original source. Browse all topics →

+ Follow this topic

Sunday, 26 January 2025

Qwen2.5-1M: Deploy Your Own Qwen with Context Length up to 1M Tokens

GitHub Pages 1 year ago 46 3 sources

Alibaba released open-source Qwen2.5-1M models, including 7B and 14B parameter versions that support context lengths up to 1 million tokens. The inference framework achieves 3.2x to 6.7x faster processing speeds on 1M-token sequences compared to baseline approaches, with the 14B model matching GPT-4o-mini performance on short texts while supporting eight times longer context. Developers can now deploy these models locally using the open-sourced vLLM-based framework, which requires 120GB VRAM for the 7B model and 320GB for the 14B model.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.