Signal Consequential Behaviour and Representational Fairness in the Validation of Synthetic Research
Summary
Companies and academic researchers increasingly substitute synthetic survey respondents generated by large language models for human samples, and they usually check these synthetic populations only through informal comparisons with human survey results. Drawing on the gap between intention and behaviour documented in behavioural science, the authors argue that such checks measure the wrong thing when decision makers commission synthetic research to anticipate what people will actually do. They propose a validation framework built on two requirements. First, every claim of validity should specify how closely it corresponds to human data: whether it predicts real behaviour, which of four diagnostics it covers (location, dispersion, the response process or structure), and whether it is tested against experimental effects. Second, validity should be reported separately for subgroups, because overall accuracy can hide the misrepresentation of the groups most affected by consequential decisions. The framework turns distributive, procedural and recognition justice into measurable quantities, requires counterfactual experiments within each persona, and is illustrated with electric vehicle charging tariffs and a reporting checklist.
Classification
Evidence 1
- Consequential Behaviour and Representational Fairness in the Validation of Synthetic Research arXiv (cs.CY) 2026-09-23 accessed 2026-09-25T10:29:27+00:00
Part of trends 0
No objects.
Directly linked issues 0
No objects.
Public id: fm-a076b9565348
