Signal What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks
Summary
This study borrows convergent and discriminant validity methods from the social sciences to examine 56 AI capability and safety benchmarks across 53 models. Model rankings on safety benchmarks grouped under the same concept often correlate only weakly with each other, suggesting the underlying concepts are not defined consistently. Conversely, capability benchmarks assigned to different concepts frequently correlate about as strongly as benchmarks sharing the same concept, undercutting claims of discriminant validity. Some benchmarks track a differently labeled concept more closely than their own, with BBQ-accuracy, for instance, moving more in step with reasoning benchmarks than with bias benchmarks. These results raise doubts about whether widely used benchmarks actually measure what they claim to measure.
Classification
Evidence 1
- What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks arXiv (cs.CY) 2026-09-08 accessed 2026-09-17T05:23:21+00:00
Part of trends 0
No objects.
Directly linked issues 0
No objects.
Public id: fm-5529584f4a3a
