TLDRocket
Sign in

Qwen

47 summarised stories about Qwen, each linking back to the original source. Browse all topics →

+ Follow this topic

Thursday, 14 November 2024

Extending the Context Length to 1M Tokens!

GitHub Pages 1 year ago 19

Alibaba's Qwen2.5-Turbo extends context length from 128k to 1 million tokens, enabling processing of approximately 10 novels or 150 hours of speech transcripts. The model achieves 93.1 on the RULER long-context benchmark, surpassing GPT-4's 91.6, and reduces time-to-first-token from 4.9 minutes to 68 seconds using sparse attention mechanisms. The pricing remains ¥0.3 per 1 million tokens, allowing the model to process 3.6 times more tokens than GPT-4o-mini at equivalent cost.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.