AI Briefing
KOSign in

Study reveals 'mode-hopping' in LM pre-training where models switch between memorization and generalization

·2026.10.01 09:00

Key point

OLMo3-32B scores dropped from 81% to 0% on an arithmetic eval between 2.17T and 2.19T tokens before recovering.

1 / 3

Details

The common assumption that language models steadily mature from pattern-matching to generalizable intelligence during pre-training is flawed. A new study demonstrates that models frequently and suddenly switch between parrot-like and intelligence-like computations, a phenomenon the authors call mode-hopping.

The Instability of Generalization

Using a toy evaluation suite, researchers tracked OLMo3 and Apertus models across pre-training checkpoints. They found that models abruptly latch onto memorized or in-context patterns instead of performing in-context learning. For example, on an "answer+1" arithmetic prompt, OLMo3-32B scored 81% at 2.17T tokens, dropped to 0% at 2.19T tokens, and recovered to 81.7% at 2.21T tokens. This instability affects various tasks, including multi-hop persona QA, out-of-context reasoning, and distinguishing truth from truthiness.

Ruling Out Optimization Noise

The study argues that mode-hopping is not merely a result of standard optimization fluctuations or evaluation metric artifacts. The phenomenon persists in both hard accuracy and soft probability margins, and it appears when plotting against both pre-training tokens and FLOPs. Furthermore, single optimization steps with varying batch sizes and learning rates did not cause the hopping, and averaging five checkpoints mitigated but did not eliminate the instability.

Practical Applications for Checkpoint Selection

The authors propose viewing mode-hopping as a capacity allocation problem where generalizable circuits compete with shallow ones. This insight allows for better checkpoint selection. In experiments, an intermediate OLMo3-32B checkpoint at 4.5T tokens outperformed a later checkpoint at 4.9T tokens after math SFT, scoring 36.3% on GPQA versus 29.8%. It also showed superior robustness to prefilling attacks (53% vs 21%). Additionally, selecting specific pre-training data can help stabilize these generalization dynamics.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.