AI Briefing
KO

Pre-training Progress Analysis 2019-2025: Data Improvements Contribute 3x More Than Model Improvements

·2026.09.09 01:10

Key point

In pre-training progress from 2019 to 2025, data improvements provided 3.24 times greater compute efficiency gains than model improvements.

Details

An analysis of pre-training progress from 2019 to 2025 revealed that data improvements yielded 3.24 times greater Compute Efficiency Gain than model improvements. Specifically, data improvements resulted in a 12.0x efficiency increase, while model improvements resulted in a 3.7x increase.

Independent Effects of Data and Models

The effects of data and model improvements were found to be largely independent and non-interacting. 88% of the variance in OLMES benchmark scores could be explained by the additive effects of model and data improvements. This indicates that the contributions of the two factors are distinct rather than complexly intertwined.

Intrinsic Value of Model Improvements

The core value of model architecture improvements lies not in simple FLOPs efficiency, but in removing constraints that enable large-scale computing. Innovations such as MoE, FlashAttention, and stability improvements (Norm, Initialization) have enabled the training of large models. Conversely, large models exhibit a characteristic where SGD filters out noise even when low-quality data is included, due to their excess capacity.

The Paradox of Large Models and Data Curation

In large models, aggressive data curation can actually degrade performance. Filtered datasets require training for dozens of epochs, yet they perform worse than large datasets with lower average quality. Frontier models tend to be overtrained by up to 100x relative to Chinchilla optimal to minimize inference compute.

Compound Annual Efficiency Growth (CEG)

The compound annual compute efficiency growth rate from 2019 to 2025 was recorded as 1.24x for models, 1.51x for data, and 1.57x combined. These figures are lower than the 3x estimate by Anson Ho et al. This discrepancy is attributed to unrealized scale-dependent improvements, the exclusion of inference optimizations, and the characteristics of the OLMES benchmark (10 easy tasks).

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.