Future Monitor 한국어

Signal New benchmark finds LLMs mediate political information reliably only under clear evidence

Summary

A paper introduces "Polistemics," a theory-grounded benchmark for evaluating how large language models mediate political information for citizens during elections. The benchmark is grounded in "Epistemic Modesty," a normative standard derived from citizens' epistemic agency, and tests models across controlled settings that vary informational properties such as clarity, noise, and consistency. The authors applied the benchmark to three state-of-the-art LLMs using real-world cases from the 2025 German and Dutch elections. They found that high aggregate performance scores mask systematic failures: models mediate information reliably when evidence is clear, but their performance breaks down when information is absent, vague, or contradictory, and they tend to flatten the intensity of political language. The paper attributes these failures in part to "party priors", biases influenced by party labels and the language in which a query is posed.

Classification

Region menusGlobal
Impactscope:global
Time horizon0-3 years (2026-07-30)
Last updated2026-07-30T08:07:57.425618+00:00

Evidence 1

Part of trends 0

No objects.

Directly linked issues 0

No objects.

Public id: fm-9401cadf0793