Signal Cultural Misalignment in Large Language Models: Detection, Measurement, and Mitigation Through Targeted Fine-Tuning
Summary
The researchers evaluated three open-weight LLMs -- Gemma3-12B from the USA, Bielik-11B-v3 from Poland, and Qwen3-4B from China -- against World Values Survey Wave 7 data for 63 demographic personas across three countries, using normalized Wasserstein distance to quantify distributional misalignment. Contrary to expectations, no model favored its home country: the Chinese-built Qwen3-4B performed worst on its own Chinese population (W1 = 0.436, the highest misalignment in the entire model-by-country matrix). Targeted LoRA fine-tuning on the five worst-case personas, requiring fewer than 1,200 training pairs and under 15 minutes on a single GPU, reduced bias by 16.8% for Bielik-11B (p_Bonf = 0.002, d = -4.4), with all five targets improving. However, country-level decomposition revealed that fine-tuning redistributed rather than removed bias: Bielik's worst-case personas swapped entirely from American to Chinese elderly, with zero overlap between the pre- and post-correction sets. The authors state this is, to their knowledge, the first study to target worst-case demographic personas with LoRA fine-tuning for cross-cultural bias mitigation.
Classification
Evidence 1
- Cultural Misalignment in Large Language Models: Detection, Measurement, and Mitigation Through Targeted Fine-Tuning arXiv (cs.CY) 2026-09-03 accessed 2026-09-17T05:23:15+00:00
Part of trends 0
No objects.
Directly linked issues 0
No objects.
Public id: fm-6b3497c3c50e
