AI Briefing
KO

The Bitter Lesson for Data Filtering

·2026.05.21 09:00

Key point

New research shows that with sufficient compute power, training large models on low-quality data without filtering is possible.

Details

There is a common belief that in large language model pretraining, selecting only high-quality data is essential. However, new research presents results suggesting that, given sufficient compute power, not filtering at all may actually be the best approach.

The research team conducted scaling experiments in a high-compute, data-scarce environment. It was found that sufficiently trained large-parameter models not only withstand low-quality or disruptive data, but actually benefit from data that is nominally "bad."

This extends Rich Sutton's "The Bitter Lesson" into the domain of data curation, suggesting that compute scale can be more effective than manual filtering.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.