Signal Can public chat data predict real-world AI misalignments?
Summary
OpenAI's alignment team examined whether public datasets like WildChat can serve as effective proxies for predicting real-world model misalignment without access to internal production data, finding that WildChat predicts production rates of undesirable behaviour within roughly 3x error on average for GPT-5.1, 5.2 and 5.4. The method, called 'deployment simulation,' was validated across four GPT-5-series Thinking model deployments from August 2025 through March 2026, analysing roughly 1.3 million de-identified conversations across 20 tracked categories spanning disallowed content and misaligned actions. In the most rigorous, outcome-blinded test on GPT-5.4 — where researchers were locked out of realized production numbers before freezing their forecast — the median multiplicative error was 1.5x, and for categories that shifted most between model versions, directional accuracy reached 92%, versus 54% for a challenging-prompts baseline. Predictive performance degrades most for technical and agentic forms of misalignment where current production use has moved furthest from WildChat's original scope. The work addresses a systematic 'evaluation gap' flagged in the 2026 International AI Safety Report, where pre-deployment results consistently fail to predict real-world model behaviour.
Classification
Evidence 1
- OpenAI Alignment Blog 2026-06-16 accessed 2026-07-28T13:59:37+00:00
Part of trends 0
No objects.
Directly linked issues 0
No objects.
Public id: fm-b1484691cf84