AI Briefing
KO

Study Finds Safety Training for GPT Models Transforms Rather Than Reduces Sexism

·2026.09.19 13:07

Key point

A study analyzing models from GPT-2 to GPT-5 identified a phenomenon of 'harm laundering,' where safety training transforms rather than eliminates sexism.

Details

A study accepted to the EMNLP 26 Main Conference analyzed 15 OpenAI models from GPT-2 to GPT-5, revealing that existing safety evaluation methods overlook a phenomenon of 'harm laundering' where sexist content is transformed rather than removed.

The research team analyzed 450,000 gender-oriented output data points and discovered the following patterns:

  • Transformation of Discriminatory Content: Clusters related to sexual violence, which were common in outputs targeting women in GPT-2, disappeared in GPT-4; however, this was found to be a transformation of form rather than a reduction in discrimination.
  • Representation Gap: Outputs targeting men gained positive representation domains such as care, emotional expression, and ally identity, whereas outputs targeting women did not.
  • GPT-5 Case: In GPT-5, a topic (Topic 5) framing breast cancer as a men's rights issue was discovered, yet three independent classifiers evaluated it as non-toxic.

Additionally, topic diversity in outputs targeting women decreased by 36% compared to men at the GPT-4 alignment boundary. The representation harm gap measured by the REGARD metric showed a positive correlation (ρ = +0.55) with model release dates, while Detoxify scores showed no significant correlation (ρ = -0.23). The research team emphasized that a decrease in toxicity scores is not a sufficient condition for a reduction in harm and proposed a three-step detection protocol applicable to all generative models.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.