Cosmopedia, a 25 Billion Token Dataset, Released
Key point
HuggingFace has released Cosmopedia, a large-scale open synthetic dataset for LLM pretraining.
Details
HuggingFace has released the Cosmopedia dataset, which helps pretrain high-performance LLMs using synthetic data, similar to Microsoft's Phi models.
Cosmopedia was generated using Mixtral-8x7B-Instruct-v0.1 and includes various types of content such as textbooks, blog posts, stories, and WikiHow articles.
- Scale: It consists of over 30 million files and 25 billion tokens, making it the largest synthetic dataset released to date.
- Transparency: Unlike existing models where the data generation process is opaque, HuggingFace has open-sourced the data generation code and the entire pipeline.
- Deliverables: Along with the dataset, HuggingFace is also releasing cosmo-1b, a 1B-scale model trained on it, together with the full pipeline code.
This release is expected to become an important resource for researchers and developers looking to train models from scratch using large-scale synthetic data.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.