Signal Moral Entropy: Auditing Bias and Uncertainty in Moral Judgment
Summary
The researchers propose Moral Entropy, a Bayesian framework that treats disagreement among annotators over moral content as information rather than noise to be voted away. It keeps a full posterior over the true label and splits its entropy into aleatoric uncertainty, meaning irreducible disagreement, and epistemic uncertainty, meaning gaps from insufficient or noisy annotation. Applied across three corpora and fifteen discourse domains, the framework shows that the permissive 'any-annotator' rule disagrees with the calibrated posterior on about 30 percent of items, with most of those disagreements being false positives. By contrast, the stricter majority and two-vote rules miss between 63 and 83 percent of true positives. The findings point to a systematic bias in how moral-judgment datasets used to train and evaluate AI systems are constructed.
Classification
Evidence 1
- Moral Entropy: Auditing Bias and Uncertainty in Moral Judgment arXiv (cs.CL/cs.CY/stat.ML) 2026-09-18 accessed 2026-09-24T01:30:12+00:00
Part of trends 0
No objects.
Directly linked issues 0
No objects.
Public id: fm-dce0ba8dfb65
