APEACH Releases Crowdsourced Hate Speech Evaluation Dataset
Key point
APEACH, a Korean hate speech evaluation dataset generated via crowdsourcing, was released to address the bias and copyright issues of existing web-crawled data.
Details
Existing Korean hate speech evaluation datasets relied on web crawling, resulting in topic bias, training data duplication, and copyright issues. To address this, APEACH introduced a method where an unspecified number of people (Crowd) directly generate sentences.
Crowdsourcing-Based Data Generation Process
APEACH consists of three stages: Pseudo Classifier, data generation and feedback, and post-labeling. Crowd workers write hateful or general sentences according to prompts, then compare the predictions of a secondary AI classifier with their own intent to correct labels. Subsequently, managers perform post-labeling to filter out critical data such as copyright violations or references to specific individuals.
Advantages Over Existing Datasets
Compared to existing datasets such as BEEP!, APEACH covers not only community colloquialisms but also written-style and polite forms of discriminatory expressions. Additionally, it exhibits less performance variance depending on the pre-training data domain (KcBERT vs SoongsilBERT), enhancing the fairness of model evaluation. Prompts were structured into 10 categories referencing the Pycon KR Code of Conduct, and the use of proper nouns was restricted for privacy protection.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.