Harm Laundering in GPT Models: Gender Discrimination Transformed Rather Than (arxiv.org)

🤖 AI Summary
A recent study has revealed a concerning phenomenon termed "harm laundering" in GPT models, where gender discrimination is not eliminated but transformed across generations. Researchers analyzed around 450,000 gender-directed outputs from models ranging from GPT-2 to GPT-5, demonstrating that explicit discriminatory content, particularly toward women, changes form rather than being reduced. For instance, while harmful depictions of women in GPT-2 outputs decrease in GPT-4, representations of men gain positive attributes, highlighting a disparity that becomes pronounced in GPT-5 outputs. This finding is significant for the AI/ML community as it challenges the reliance on toxicity scores as indicators of improved model safety. The research indicates that a reduction in toxicity ratings does not equate to a real decrease in harm, suggesting a need for more nuanced evaluation methodologies. By formalizing harm laundering with a three-criteria test and a detection protocol for generative models, the study calls attention to the complexity of biases in AI systems and underscores the importance of developing more effective mechanisms to assess and mitigate discrimination in AI outputs.
Loading comments...
loading comments...