Every AI story that matters — and the intelligence behind it.
TLDRocket reads all relevant sources, removes duplicate coverage, and publishes a short neutral
summary of every story, linking back to the original. Free, no spam, unsubscribe anytime.
Alibaba's Qwen released Qwen2.5, a family of open-source language models in sizes from 0.5B to 72B parameters, along with specialized variants for coding and mathematics. The 72B model achieves MMLU scores above 85, HumanEval coding scores above 85, and MATH scores above 80, trained on 18 trillion tokens. The models support 128K token context windows, multilingual support for 29+ languages, and improved capabilities in long text generation, structured data understanding, and JSON output generation.
Alibaba released Qwen2.5, a series of seven open-source language models ranging from 0.5B to 72B parameters, with new mid-size models at 14B and 32B designed for production use. The pre-training dataset expanded from 7 trillion to 18 trillion tokens, and Qwen2.5-72B achieved an MMLU score of 86.1 compared to Qwen2-72B's 84.2, while Qwen2.5-32B outperformed the larger Qwen2-72B in various benchmarks. The models show significant improvements in coding, mathematics, and long-context generation, with Qwen2.5-72B-Instruct reaching an Arena-Hard score of 81.2 and LiveCodeBench score of 55.5.
Alibaba released Qwen2.5-Coder, an open-source coding model family trained on 5.5 trillion tokens of code and other data. The 7B version outperforms larger models like DeepSeek-Coder-V2-Lite and CodeStral-22B on code benchmarks, while maintaining math and general knowledge capabilities similar to the base Qwen2.5 model. A 32B version is forthcoming as the company aims to compete with proprietary code models.
Alibaba released Qwen2.5-Math, an updated series of open-source mathematical language models in sizes 1.5B, 7B, and 72B parameters. The 72B-Instruct model achieved a score of 92.9 on the MATH benchmark using tool-integrated reasoning and solved 12 problems on the AIME 2024 exam compared to 1-2 problems for GPT-4 and Gemini models. The models now support both English and Chinese math problems using chain-of-thought and tool-integrated reasoning approaches, with the 7B model matching previous 72B model performance.
Researchers introduced CORE-Bench, a benchmark for measuring how well AI agents can automate computational reproducibility in scientific research. The best AI agent tested (CORE-Agent with GPT-4o) achieved 22% accuracy on the hardest difficulty level despite task-specific modifications. The work suggests that AI systems may prove economically valuable for automating specific scientific tasks even without general-purpose capabilities, challenging conventional notions of artificial general intelligence.
Every AI story that matters,
in your inbox by 8am.
TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the
day in two minutes. Follow companies and topics for alerts, or get the
briefing in Slack. Free, no spam, unsubscribe anytime.
Reading TLDRocket needs no cookies, and the readership counts we rely on come from
our own cookieless analytics. Google Analytics is the exception: it sets cookies and
reports to Google, so it stays switched off until you allow it. You can change your
mind any time from “Cookie settings” in the footer.
Strictly necessary
Session security and form protection (tldrocket-session,
XSRF-TOKEN, 2 hours). The site cannot work without them,
so they need no consent.
Always on
Google Analytics 4 (_ga,
_ga_<id>, up to 2 years). Measures which
stories and sections readers use. Google acts as a third-party processor and may
store the data outside the EU. No advertising, no profiling, no data sold.