Extending the Context Length to 1M Tokens!
Qwen
Alibaba's Qwen2.5-Turbo extends context length from 128k to 1 million tokens, enabling processing of approximately 10 novels or 150 hours of speech transcripts. The model achieves 93.1 on the RULER long-context benchmark, surpassing GPT-4's 91.6, and reduces time-to-first-token from 4.9 minutes to 68 seconds using sparse attention mechanisms. The pricing remains ¥0.3 per 1 million tokens, allowing the model to process 3.6 times more tokens than GPT-4o-mini at equivalent cost.
Why it matters
API Documentation (Chinese) HuggingFace Demo ModelScope Demo Introduction After the release of Qwen2.5, we heard the community’s demand for processing longer contexts. In recent months, we have made many optimizations for the model capabilities and inference performance of extremely long context. Today, we are proud to introduce the new Qwen2.5-Turbo version, which features: Longer Context Support: We have extended the model’s context length from 128k to 1M, which is approximately 1 million English words or 1.