Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations
July 29, 2026
Fairness Pruning localizes demographic bias in GLU architectures by identifying neurons with differential activations at the down_proj input using minimally contrastive prompt pairs. Testing on Llama-3.2 and Salamandra-2B models shows that zeroing these specific neurons enables targeted structural intervention to mitigate bias without full retraining.