Future Monitor 한국어

Signal Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations

Summary

Researchers presented Fairness Pruning, a lightweight structural intervention method for locating and eventually mitigating demographic bias in large language models, with this work focused on causal bias localization. Using minimally contrastive prompt pairs and inference-time activation capture, the method identifies neurons in GLU-MLP architectures that react differentially when models process demographic attributes. The method was empirically evaluated on models up to 3 billion parameters, including the Llama-3.2 family and Salamandra-2B, combining benchmark evaluation with qualitative text-generation experiments. Results showed that zeroing the identified neurons changes how models respond to associated demographic variables, but rather than producing uniform mitigation, the intervention causes what the authors call 'bidirectional bias destabilization,' since the identified neuron sets mix those pushing toward and against stereotypes. The intervention was found to be highly targeted: zeroing at most 40 neurons in Llama-3.2-1B, less than 0.031% of total MLP width, achieved 99.49% mean retention of reasoning and general knowledge capabilities. The authors state these findings confirm that demographic bias processing and model capabilities operate on separable circuits within the model.

Classification

Main topicAI & Computing
Region menusGlobal
Impactscope:global
Time horizon4-10 years (2026-08-01)
Last updated2026-08-01T03:55:37.548876+00:00

Evidence 1

Part of trends 0

No objects.

Directly linked issues 0

No objects.

Public id: fm-3b6ce2b11c73