AI article
Harm Laundering in GPT Models: Gender Discrimination Transformed Rather Than
No preview is available. Read the original article for the full story.
Hacker News | Sep 19, 2026 | sbulaev
Automated excerpt
Abstract:Safety evaluations for large language models rely on surface-form classifiers that report declining harm scores across model generations. Analysing 450,000 gender-directed completions across 15 models spanning GPT-2 through to GPT-5 (OpenAI GPT lineage; three demographic conditions), we show that sexual violence clusters prevalent in GPT-2 women-directed output disappear by GPT-4, while men-directed completions gain positive representational territory (caregiving, emotional range, ally identity) that women-directed completions do not. Sentiment scores invert at GPT-4: early models demean women; later models over-correct.
Selected automatically from source text; not independently written or fact-checked. Read the original for full context.