[HN]score: 0.24
Harm Laundering in GPT Models: Gender Discrimination Transformed Rather Than
September 19, 2026
Analysis of 450,000 completions across GPT-2 through GPT-5 reveals that safety training transforms explicit gender discrimination into non-toxic representational harm rather than reducing it. While toxicity scores decline, topic diversity for women-directed output drops 36% relative to men at the GPT-4 boundary, with representational harm correlating positively with model release date.
DAILY DIGEST
you don't check 9 sources — we do. one email every morning, read in 2 min. free. unsubscribe anytime. privacy