AI Briefing
KO

Measuring LLM Memorization Capacity and Scaling Laws

·2026.07.07 18:32

Key point

The paper proposes a new methodology that separates memorization and generalization in language models to measure model capacity.

Details

To address the difficulty prior research has had in distinguishing between Memorization and Generalization in language models, this paper proposes a new measurement method that formally separates the two concepts.

According to the research, a model's total memorization capacity is found to be approximately 3.6 bits per parameter for GPT-style models. The researchers trained hundreds of Transformer models ranging from 500K to 1.5B parameters and observed the following phenomena.

  • Capacity Saturation and Grokking: Once a model fills up its capacity with training data, the 'Grokking' phenomenon begins, during which unintended memorization decreases and generalization starts.
  • Scaling Laws: The researchers derived a set of scaling laws representing the relationship between model capacity, data size, and membership inference.

This research provides a framework for quantitatively distinguishing between the process by which a model learns the generative process behind the data and the process by which it simply memorizes.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.