Future Monitor 한국어

Signal Study finds LLMs can approximate aggregate housing-survey results but fail to preserve population structure

Summary

Researchers tested whether large language models can serve as low-cost proxies for resident attitudes in urban planning by comparing eight open-weight LLMs against 843 human respondents in a US affordable-housing survey experiment. The study examined whether models could reproduce the difference in support change between homeowners and renters as a proposed housing development moved from two miles to one-eighth of a mile away. Qwen 2.5 14B came closest to the human results (-0.242 versus human -0.285) and was the only model meeting the prespecified equivalence criterion, while other models showed weak, null or reversed effects. However, this aggregate match masked structural failures: Qwen attenuated the Republican-leaning contrast while exaggerating the Independent-leaning one, showed a root-mean-square error of 0.613 across 27 party-by-tenure-by-item cells, and had a median model-to-human variance ratio of just 0.099. Question order alone shifted the measured contrast by 0.367, and removing identity cues or accounting for selective nonresponse changed which comparisons were even estimable. The researchers conclude an LLM can approximate one aggregate effect while failing to preserve the underlying population structure and measurement stability, and argue model evaluation in urban planning should test whether this structure survives simulation, not just average effects.

Classification

Region menusGlobal
Impactscope:global
Time horizon4-10 years (2026-07-31)
Last updated2026-07-31T01:34:05.815397+00:00

Evidence 1

Part of trends 0

No objects.

Directly linked issues 0

No objects.

Public id: fm-315ed7781174