Start your day with intelligence. Get The OODA Daily Pulse.

Home > Briefs > Technology > Pretraining progress is mostly coming from data

Pretraining progress is mostly coming from data

How much of the rapid progress in AI that we’ve seen over the last few years1 has come from data versus model improvements? The answer has big implications for the economics of frontier labs and the pace of future progress. We investigate this question at a relatively small scale, and for pretraining specifically, from 2019 to 2025. During each of those years, a new open model recipe was published which codified that year’s publicly known algorithmic tweaks (for example, improvements in architecture, optimizer, initializations, learning rate schedule, hyperparams, etc). And during each of those years, there was also a new public data corpus (produced by broader scrapes and new curation/extraction/filtering techniques).

Full opinion : Analysis from 2019 to 2025, gains in pretraining compute efficiency came mostly from data improvements rather than model improvements.