A public dashboard observing signals, trends and issues.
SubscribeLogin한국어
Latest observation
2026-10-08
Public objects
4434
Build time
2026-10-08 19:44 KST
The Futures

Signal What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks

Summary

This study borrows convergent and discriminant validity methods from the social sciences to examine 56 AI capability and safety benchmarks across 53 models. Model rankings on safety benchmarks grouped under the same concept often correlate only weakly with each other, suggesting the underlying concepts are not defined consistently. Conversely, capability benchmarks assigned to different concepts frequently correlate about as strongly as benchmarks sharing the same concept, undercutting claims of discriminant validity. Some benchmarks track a differently labeled concept more closely than their own, with BBQ-accuracy, for instance, moving more in step with reasoning benchmarks than with bias benchmarks. These results raise doubts about whether widely used benchmarks actually measure what they claim to measure.

Classification

Main topicAI & Computing
Region menusGlobal
Impactscope:global
Time horizon0-3 years (2026-09-17)
Last updated2026-09-25 22:32 KST

Evidence 1

Part of trends 0

No objects.

Directly linked issues 0

No objects.

Public id: fm-5529584f4a3a