AI article

Harm Laundering in GPT Models: Gender Discrimination Transformed Rather Than

No preview is available. Read the original article for the full story.

Hacker News | Sep 19, 2026 | sbulaev

Automated excerpt

Abstract:Safety evaluations for large language models rely on surface-form classifiers that report declining harm scores across model generations. Analysing 450,000 gender-directed completions across 15 models spanning GPT-2 through to GPT-5 (OpenAI GPT lineage; three demographic conditions), we show that sexual violence clusters prevalent in GPT-2 women-directed output disappear by GPT-4, while men-directed completions gain positive representational territory (caregiving, emotional range, ally identity) that women-directed completions do not. Sentiment scores invert at GPT-4: early models demean women; later models over-correct.

Selected automatically from source text; not independently written or fact-checked. Read the original for full context.

Read the original article

More AI news