TLDRocket
Sign in

Pretraining Progress Is Mostly Coming From Data

Dwarkesh Podcast

Researchers compared pretraining recipes and data corpora from 2019 to 2025 and found more of the compute-efficiency gains came from data than from model changes. Data improvements contributed 3.24x more compute efficiency gains than model improvements (12.0x for data vs 3.7x for models) at a 1e19 FLOPs budget. The study concludes data and model improvements are mostly independent additively (88% of OLMES score variance explained), so progress is increasingly driven by better data engineering rather than interacting with specific model tweaks.

Why it matters

The issue highlights a controlled study examining open model recipes and datasets from 2019 to (the rest of the article summary isn’t included in the provided text).

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.