AI Briefing
KO

LittleLearner (Website)

·2026.08.17 09:00

Key point

LittleLearner demonstrated that pre-training data determines the knowledge acquisition limits of models by training on specific educational curriculum data.

Details

Modern language models (LMs) are trained on vast amounts of data simultaneously, making it difficult to distinguish whether a specific skill has been truly 'learned' by the model or simply 'elicited'. LittleLearner established a controlled sandbox environment using LittleCurriculum (88B tokens), filtered to align with the US elementary school (K-5) curriculum, to study this phenomenon.

LittleLearner is available in three sizes: 0.6B, 1.3B, and 5B. Each model is accompanied by an Unfiltered control group that shares the same architecture and recipe. The models consist of pre-trained Base, math-specialized GRPO, and general chat Chatty variants.

Experimental results showed that model scaling (Scaling), post-training via GRPO (Post-training), and in-context learning (In-context learning) all improved performance within the curriculum scope (In-scope) but failed to improve performance outside the scope (Out-of-scope). This demonstrates that data filters during the pre-training stage determine the model's actual capability ceiling.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.