AI Briefing
KO

Hugging Face Scans 7.6 Petabytes of Training Data for Secrets

·2026.06.01 09:00

Key point

220,000 valid credentials were found in Hugging Face's public training data.

Details

Truffle Security announced that after scanning all of Hugging Face's public datasets—7.6PB, 187 million files—it discovered 221,303 valid, non-duplicate credentials across 6,003 datasets.

One of the credentials with potentially the largest impact had access to approximately 393GB of personally identifiable information (PII), estimated to correspond to about 3.7% of the world's population. Specific details will be disclosed in a follow-up report.

Write-access tokens that could lead to supply chain attacks were also identified.

  • 349 GitHub Personal Access Tokens: 223 with full repo write access, 130 with CI workflow modification rights, 112 with admin:org permissions, 110 with package publishing rights
  • 318 Docker Hub push tokens
  • 237 Hugging Face write tokens and 70 organization admin tokens

Some tokens were linked to the account of the founder of a widely used MCP registry, and the repositories under that organization—which include servers and SDKs used by major AI coding tools—were found to have a combined total of over 178,000 GitHub stars. The researchers stated that they did not disclose the names of the individuals or organizations involved and notified them responsibly.

Infrastructure access was also extensive.

  • 8,557 GCP service account keys: linked to 3,811 projects
  • Access to 51.7TB of private S3 buckets
  • 8,594 working database logins, totaling 3.5TB based on metadata
  • 5,885 Slack and Mailgun keys

The researchers stated that for verification and impact assessment, they only checked metadata such as database size, Redis memory usage, and S3 bucket capacity, and did not read database rows or object listings, download files, or modify any systems. Hugging Face cooperated with the investigation, and CTO Julien Chaumond added storage bucket scanning functionality to TruffleHog.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.