Signal Follow the Norm: Accounting for Fine-Tuning and Prompt Effects on Model Rationales
Summary
Normative datasets commonly used to train and align AI systems can function as action-guiding patterns rather than neutral moral knowledge. Treating an AI system as a proxy actor, this paper tests whether dataset-level norms introduced through fine-tuning and prompting can push a system away from its baseline safety behavior when facing high-conflict moral dilemmas. The authors present three contributions showing that both the composition of fine-tuning data and the framing of prompts have measurable effects on the rationales a model produces. This demonstrates that training data which appears neutral on its surface can still covertly steer model behavior in ethically fraught situations. As a result, the work argues that normative datasets themselves deserve scrutiny as part of AI safety and alignment verification.
Classification
Evidence 1
- Follow the Norm: Accounting for Fine-Tuning and Prompt Effects on Model Rationales arXiv (cs.CY) 2026-08-13 accessed 2026-08-16T10:59:38+00:00
Part of trends 0
No objects.
Directly linked issues 0
No objects.
Public id: fm-d51c403b6681
