Signal Difference-in-Differences on a Censored Rating Scale Can Manufacture an Effect: Evidence from a Pre-Registered LLM-Judge Audit
Summary
Audits that use an LLM as a judge to check for bias typically rely on a strong difference-in-differences design that contrasts conditions within the same item. This paper carries out an actual pre-registered LLM-judge audit and shows that this design can cause problems in certain situations. Specifically, when the design is applied to a censored rating scale, it can manufacture an effect that does not actually exist. This suggests that some existing bias-audit findings produced this way could be spurious. The material available here is limited to the abstract-level claim, and the specific sample size and quantitative results can only be confirmed by reading the full paper.
Classification
Main topicAI & Computing
Region menusGlobal
Impactscope:global
Time horizon0-3 years (2026-08-29)
Last updated2026-09-25 22:32 KST
Evidence 1
- Difference-in-Differences on a Censored Rating Scale Can Manufacture an Effect: Evidence from a Pre-Registered LLM-Judge Audit arXiv (cs.CY) 2026-08-27 accessed 2026-09-08T08:09:20+00:00
Part of trends 0
No objects.
Directly linked issues 0
No objects.
Public id: fm-7ac244e41944
