Unbounded Labs Releases Vintage LLM 'Bart'
Key point
Unbounded Labs has open-sourced 'Bart', a 2.82B parameter LLM trained on text from before 1931.
Details
Unbounded Labs has released 'Bart', an LLM with 2.82B parameters trained on 20.1B tokens of English text from before 1931. The model was developed to explore whether LLMs can reproduce conclusions reached by scientists of the past, inspired by a suggestion from Demis Hassabis.
Key achievements and released materials include:
- Dataset: Curated Harvard Institutional Books (242B tokens) to create a high-quality corpus of 23B tokens.
- Benchmarks: Developed the first evaluation suite specifically for vintage LLMs, Vintage CORE (20 benchmarks).
- Model Performance: Achieved higher performance than GPT-1900 at this scale with a smaller token budget.
- Open Source: Released all assets, including an SFT dataset consisting of 416,000 Q&A pairs, training code, evaluation metrics, and training logs.
The model was completed after 5 days of training on H100 GPUs and is available on Hugging Face, accompanied by a blog post that details the entire research process and mistakes.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.