AI Briefing
KO

Hugging Face experiments with PII detection for datasets

·2024.07.10 09:00

Key point

Hugging Face is experimenting with a report feature that uses Presidio to flag whether datasets contain personally identifiable information (PII).

Details

Hugging Face is experimenting with a new feature to address the problem of undisclosed personally identifiable information (PII) being found in machine learning datasets uploaded to the Dataset Hub.

Key points are as follows:

  • Problem recognition: Large-scale, web-crawled pretraining datasets can leak subtle personal information, leading to privacy violations and model bias.
  • Solution: Using the open-source PII detection tool Presidio, the feature provides a report estimating whether PII is present in a dataset.
  • Expected benefits:
    • Users: Can check whether a dataset contains sensitive information before training and decide whether additional filtering is needed.
    • Dataset owners: Can validate the effectiveness of their PII filtering process before releasing data.

This feature is currently in an experimental stage, and detection results for sensitive information such as emails can be viewed via the allenai/c4 dataset example.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.