Alignment Reduces Expressed but Not Encoded Gender Bias in LLMs
September 11, 2026
A unified framework reveals that while alignment techniques reduce gender bias in model outputs, the underlying latent representations remain biased. This finding indicates that alignment masks rather than eliminates internal social regularities learned during training.
HOW THIS AFFECTS YOU
●
researcherYou should investigate latent representations rather than relying solely on output-level benchmarks for safety evaluation.
●
policyBe aware that superficial alignment may hide persistent biases within the model's core architecture.