Understanding Data Annotator Safety Policies Through Interpretability
Key point
This work proposes APM models that learn internal safety policies from data annotators' labeling behavior and analyze the causes of disagreement.
Details
Disagreements that arise during the data annotation process, which is central to AI model development, stem from various causes such as operational failure, policy ambiguity, or value pluralism. However, directly asking annotators for their reasons is costly, and there is a limitation in that the self-reported reasons may differ from the actual decision-making process.
To address this, Annotator Policy Models (APMs) are introduced. Without requiring additional explanation requests, APMs learn solely from labeling behavior data to model annotators' internal safety policies in an interpretable form. This model reproduces annotators' policies with over 80% accuracy and shows reliable predictive performance even for counterfactual edits.
APMs offer two main core applications:
- Identifying policy ambiguity: Revealing how annotators interpret safety guidelines differently from one another
- Detecting value pluralism: Systematically identifying differences in safety priorities across demographic groups
This technology can contribute to designing AI safety policies that are more sophisticated, transparent, and capable of embracing diverse perspectives.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.