TLDRocket
Sign in

Qwen2.5-LLM: Extending the boundary of LLMs

Qwen Covered by 3 sources

Alibaba released Qwen2.5, a series of seven open-source language models ranging from 0.5B to 72B parameters, with new mid-size models at 14B and 32B designed for production use. The pre-training dataset expanded from 7 trillion to 18 trillion tokens, and Qwen2.5-72B achieved an MMLU score of 86.1 compared to Qwen2-72B's 84.2, while Qwen2.5-32B outperformed the larger Qwen2-72B in various benchmarks. The models show significant improvements in coding, mathematics, and long-context generation, with Qwen2.5-72B-Instruct reaching an Arena-Hard score of 81.2 and LiveCodeBench score of 55.5.

Why it matters

GITHUB HUGGING FACE MODELSCOPE DEMO DISCORD Introduction In this blog, we delve into the details of our latest Qwen2.5 series language models. We have developed a range of decoder-only dense models, with seven of them open-sourced, spanning from 0.5B to 72B parameters. Our research indicates a significant interest among users in models within the 10-30B range for production use, as well as 3B models for mobile applications. To meet these demands, we are open-sourcing Qwen2.

Also covered by

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.