AI Briefing
KO

103B-token Usenet corpus released

·2026.05.02 03:01

Key point

A 33-year, 103.1B-token Usenet corpus and its cleaning methodology have been released.

Details

A 103.1B-token corpus compiling the entire Usenet archive from 1980–2013 has been organized. Its scale spans 408 million posts, 18,347 newsgroups, and 33 consecutive years of coverage.

The cleaning process included the following:

  • Deduplication
  • Excluding the alt.binaries.* hierarchy
  • Quote handling
  • Masking email addresses
  • Hashing Message-IDs with SHA-256
  • Converting the original MBOX files into gzip-compressed JSONL

Applying Meta's fastText LID-176 to every record showed that 96.6% of the corpus was in English, with over 100 languages mixed in. In particular, the soc.culture.* groups have a high proportion of non-English content.

Over time, volume was sparse before 1986, grew rapidly from the early 1990s, peaked around 1999–2000, and then declined as forums and social media took over. The datacard, cleaning methodology, and representative samples have been released on Hugging Face.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.