Signal LLM Alignment Should Go Beyond Harmlessness-Helpfulness and Incorporate Human Agency
Summary
The paper, titled "LLM Alignment should go beyond Harmlessness-Helpfulness and incorporate Human Agency," was published in the journal Cognitive Computation (Springer Nature, 2026). It argues alignment is inherently multidimensional, spanning safety, helpfulness, fairness and justice, and cultural sensitivity, rather than reducible to a harmlessness-versus-helpfulness tradeoff. The authors propose shifting from static, post-training constraints toward dynamic, participatory approaches that safeguard pluralism, autonomy, and human flourishing, formalized as a Flourishing-Justice-Autonomy (FJA) framework. They contrast this with conventional Harmless-Helpful-Honest (HHH) aligned models, which use rigid post-training interventions and often produce overcautious outputs, versus FJA-aligned models, which would employ inference-time adaptability, participatory constitutions, and dynamic reward models. The framing shifts alignment from output filtering toward preserving the user's capacity to decide. The specific article page could not be directly accessed due to a login redirect loop in an earlier attempt, so author names and exact issue/volume details were not independently verified beyond the journal and title identified via search.
Classification
Evidence 1
- Springer Nature / arXiv / NC State 2026-01-01 accessed 2026-07-28T13:59:43+00:00
Part of trends 0
No objects.
Directly linked issues 0
No objects.
Public id: fm-dc3153169d1a