AI Briefing
KO

Hugging Face Releases The Stack v3

·2026.07.24 20:57

Key point

Hugging Face has released The Stack v3, the largest open-source code dataset to date.

Details

Hugging Face has released The Stack v3, the largest open code dataset to date. This release comes in two forms depending on the user's purpose.

  • stack-v3-train: A near-deduplicated, quality-filtered, and PII-removed dataset, structured to be ready for training right away.
  • stack-v3-full: The entire 114TB corpus provided as an HF Storage Bucket, retaining all duplicate data along with cluster IDs so that users can perform their own filtering and mixing.

Developers can load the dataset directly for code model training and research.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.