Fairness Pruning Locates Demographic Bias in GLU-MLP Layers
July 31, 2026
Fairness Pruning identifies biased neurons in GLU architectures by capturing differential activations during demographic attribute processing. Testing on Llama-3.2 and Salamandra-2B shows that zeroing these neurons alters demographic responses, though it can lead to bidirectional bias shifts.
HOW THIS AFFECTS YOU
●
researcherYou can use this structural intervention to localize and mitigate causal bias in LLM layers.
●
policyThis provides a technical method for auditing and addressing demographic bias in model architectures.