A public dashboard observing signals, trends and issues.
SubscribeLogin한국어
Latest observation
2026-10-08
Public objects
4434
Build time
2026-10-08 19:44 KST
The Futures

Signal New Benchmark Isolates Why Video Language Models Fail at Basic Event Counting

Summary

The preprint notes that real-world video benchmarks offer broad coverage, but their fixed clips mix together event count, rate, duration, and visual complexity, making it hard to pinpoint why a model fails. Existing programmatically generated benchmarks control these factors better, but they only score the final answer rather than checking reported events against executable ground truth. To close this gap, the authors introduce a trace-grounded parametric profiling benchmark that can systematically vary these factors and verify reported events against checkable ground truth. Experiments using this benchmark show that once confounding factors are properly separated, video language models struggle even with the basic task of accurately counting or tracking distinct events within a video. This finding suggests that reported performance on entangled real-world benchmarks may have masked a more fundamental weakness. The result points to event counting as a meaningful diagnostic for evaluating video understanding models.

Classification

Main topicAI & Computing
Region menusGlobal
Impactscope:global
Time horizon0-3 years (2026-08-10)
Last updated2026-09-25 22:32 KST

Evidence 1

Part of trends 0

No objects.

Directly linked issues 0

No objects.

Public id: fm-bad84b7489da