Signal Participatory Moral AI Is Not Neutral: The Invisible Hand of Developers
Summary
This paper examines moral preference elicitation, an approach in which researchers poll participants on hypothetical dilemmas and use the aggregated votes to train an AI policy applied at scale. The authors argue that before any vote is cast, developers make three often-opaque choices, deciding which features to scope for a vote, which voters to sample, and how to frame the question. To test this, they ran an empirical study across two phases with 809 participants in three deployment contexts, AI kidney allocation, AI agents simulating absent workers, and generative AI depictions of the deceased. They found that morally relevant features shift across contexts, and that preferences differ by political ideology for roughly a third of features, with some differences reversing direction depending on the voter pool's ideological makeup. They also found that the wording of the elicitation question can narrow or widen ideological gaps by up to a full scale point. The authors conclude that voting-based alignment alone cannot guarantee fair or transparent AI, and they recommend auditing and disclosing every stage of the moral AI elicitation pipeline.
Classification
Evidence 1
- Participatory Moral AI Is Not Neutral: The Invisible Hand of Developers arXiv (cs.AI) 2026-08-14 accessed 2026-08-20T05:08:17+00:00
Part of trends 0
No objects.
Directly linked issues 0
No objects.
Public id: fm-9b33c2a841ea
