AI Briefing
KO

HuggingFace Releases SmolLM

·2024.07.16 09:00

Key point

HuggingFace released the SmolLM series of small language models in 135M, 360M, and 1.7B sizes, trained on high-quality data.

1 / 2

Details

HuggingFace announced SmolLM, a series of state-of-the-art small language models (SLMs) designed to run efficiently on local devices. SmolLM is available in three sizes: 135M, 360M, and 1.7B parameters.

The key to these models lies in SmolLM-Corpus, a meticulously curated high-quality training dataset. This dataset consists of:

  • Cosmopedia v2: a 28B-token synthetic dataset of textbooks and stories generated via Mixtral
  • Python-Edu: 4B tokens of educational Python samples extracted from The Stack
  • FineWeb-Edu: 220B tokens of educational web samples extracted from FineWeb

SmolLM demonstrates performance that overwhelms other models of the same size class on common-sense reasoning and world knowledge benchmarks. HuggingFace open-sourced not only the models but also the entire dataset used for training, emphasizing the importance of data curation in boosting the performance of small models.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.