TLDRocket
Sign in

Qwen

47 summarised stories about Qwen, each linking back to the original source. Browse all topics →

+ Follow this topic

Thursday, 6 June 2024

Generalizing an LLM from 8k to 1M Context using Qwen-Agent

GitHub Pages 2 years ago 35

Alibaba's Qwen team built a multi-level agent system that extends an 8k-token context model to handle 1-million-token documents by combining retrieval-augmented generation, chunk-by-chunk reading, and step-by-step reasoning rather than relying on native long-context models. The system was evaluated on NeedleBench and LV-Eval benchmarks designed for 256k-context tasks, where the 4k-Agent consistently outperformed both a 32k-context model extended via RoPE extrapolation and basic RAG approaches. The agent framework is being released as open-source infrastructure to generate synthetic fine-tuning data for training long-context models.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.