Pretraining Progress Is Mostly Coming From Data
Dwarkesh Podcast
Researchers compared pretraining recipes and data corpora from 2019 to 2025 and found more of the compute-efficiency gains came from data than from model changes. Data improvements contributed 3.24x more compute efficiency gains than model improvements (12.0x for data vs 3.7x for models) at a 1e19 FLOPs budget. The study concludes data and model improvements are mostly independent additively (88% of OLMES score variance explained), so progress is increasingly driven by better data engineering rather than interacting with specific model tweaks.
Why it matters
The issue highlights a controlled study examining open model recipes and datasets from 2019 to (the rest of the article summary isn’t included in the provided text).