AI Briefing
KO

103B-Token Usenet Dataset Released

·2026.05.28 05:23

Key point

An AI-contamination-free dataset of 103B tokens built from Usenet data spanning 1980 to 2013 has been released.

Details

A 103.1B-token dataset built from the Usenet corpus spanning 1980 to 2013 has been released. The core point of this dataset is that it consists of pure human-written data from long before the emergence of LLMs, allowing it to avoid the AI contamination, GPT-specific phrasing, and RLHF bias that characterize modern web data.

The key characteristics of the dataset are as follows:

  • Zero AI contamination: All posts were written before the LLM era, so they contain no model refusal patterns or RLHF artifacts.
  • Non-algorithmic data: Composed mainly of long, substantive writing from before content was optimized for SEO or clickbait.
  • Diverse domain hierarchies:
    • comp.*: Computing-related (10.3B tokens)
    • sci.*: Science-related (3.3B tokens)
    • rec.*: Hobbies, sports, arts, etc. (16.5B tokens)
    • humanities.*: Philosophy, literature, etc.

The dataset totals 103.1B tokens (based on cl100k_base) and includes about 408 million posts. Preprocessing such as deduplication, removal of email addresses, and removal of binaries has been completed. Sample data is currently available for free download, and the full corpus is available through a license agreement.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.