AI Briefing
KO

13B Model Talkie Trained Only on Pre-1931 Data

·2026.04.29 01:06

Key point

Talkie, a 13B-scale 'vintage' language model trained solely on historical texts from before 1931, has been released, enabling contamination-free testing of generalization ability.

Details

Talkie is a 13B-scale language model trained using only text data from before 1931. As a 'vintage' model excluding modern knowledge or values, its purpose is to research AI's scope of knowledge and predictive capabilities.

Key Features and Research Content:

  • Future Prediction and Knowledge Analysis: By measuring the model's 'surprisingness' regarding historical events after 1931, it analyzes how difficult the model finds it to predict events beyond its knowledge cutoff.
  • Solving Data Contamination: Since no modern web data or code is included, it enables pure generalization experiments, such as testing whether the model can learn a modern language like Python despite having no knowledge of digital computers.
  • Coding Ability Experiments: Although its performance is lower than modern models, it shows a pattern of gradually improving coding ability as model scale increases, such as implementing simple logic through in-context learning.

This model is expected to become an important research tool for understanding changes in predictive performance according to model size and scaling laws.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.