AI Briefing
KO

Unbounded Labs Releases Vintage LLM 'Bart'

·2026.08.25 01:14

Key point

Unbounded Labs has released 'Bart', a 2.82B parameter LLM trained on data from before 1931.

Details

Unbounded Labs has launched 'Bart', an LLM with 2.82B parameters trained on 20.1B tokens of English text from before 1931. The model was developed to explore whether it can replicate the thought processes of historical scientists and was built in-house for a total cost of $807.

Key Achievements and Open Source

  • Dataset: Curated Harvard Institutional Books (242B tokens) to build a high-quality corpus of 23B tokens.
  • Benchmark: Created Vintage CORE, an evaluation suite specifically for vintage LLMs, which did not previously exist.
  • Performance: Outperformed GPT-1900 on the Vintage CORE benchmark among models of the same size.
  • Data Release: Released an SFT dataset of 416k examples based on text from before the 1930s.
  • Training: Trained the final model in 5 days on H100 GPUs, maintaining 60% MFU.

All datasets, methodologies, training code, and evaluation results are open-sourced, allowing researchers to verify and utilize them directly.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.