Signal New diagnostic framework finds LLM personas fail to update beliefs like humans in deliberative polling
Summary
Researchers introduce a Deliberative Polling Diagnostic Framework that compares human and LLM belief shifts after identical informational interventions, addressing whether LLM personas used in 'silicon sampling' to simulate public opinion actually update beliefs like humans do during deliberation rather than simply retrieving cached opinions. Applying the framework to five frontier models using data from America in One Room (526 personas, 72 questions), the authors find every model fails in a distinct way. GPT-5.1 exhibits reversal, becoming more hostile toward the opposing party after balanced information while humans become less so, a pattern that is selective (80% on outgroup versus 26% on policy questions). Gemini 2.0 Flash, Claude Sonnet 4.5, and Llama 3.3 70B exhibit overshoot, shifting in the correct direction but at 5-7 times human magnitude, while DeepSeek V3 shows near-zero change. The authors term this signature 'self-sycophancy' -- conformity to the model's internal stereotype of the persona rather than reasoning from the information provided -- and propose running the diagnostic before trusting LLM personas to mimic revised beliefs.
Classification
Evidence 1
- Before You Poll with LLMs: A Deliberative Diagnostic Framework arXiv (cs.CY) 2026-09-14 accessed 2026-09-17T05:23:29+00:00
Part of trends 0
No objects.
Directly linked issues 0
No objects.
Public id: fm-3e44fc21a235
