Future Monitor 한국어

Signal Can public chat data predict real-world AI misalignments?

Summary

OpenAI's alignment team examined whether public datasets like WildChat can serve as effective proxies for predicting real-world model misalignment without access to internal production data, finding that WildChat predicts production rates of undesirable behaviour within roughly 3x error on average for GPT-5.1, 5.2 and 5.4. The method, called 'deployment simulation,' was validated across four GPT-5-series Thinking model deployments from August 2025 through March 2026, analysing roughly 1.3 million de-identified conversations across 20 tracked categories spanning disallowed content and misaligned actions. In the most rigorous, outcome-blinded test on GPT-5.4 — where researchers were locked out of realized production numbers before freezing their forecast — the median multiplicative error was 1.5x, and for categories that shifted most between model versions, directional accuracy reached 92%, versus 54% for a challenging-prompts baseline. Predictive performance degrades most for technical and agentic forms of misalignment where current production use has moved furthest from WildChat's original scope. The work addresses a systematic 'evaluation gap' flagged in the 2026 International AI Safety Report, where pre-deployment results consistently fail to predict real-world model behaviour.

Classification

Main topicAI & Computing
Region menusGlobal
Impactscope:global
Time horizon0-3 years (2026-07-29)
Last updated2026-07-28T14:11:55.790559+00:00

Evidence 1

Part of trends 0

No objects.

Directly linked issues 0

No objects.

Public id: fm-b1484691cf84