Signal New Benchmark Isolates Why Video Language Models Fail at Basic Event Counting
Summary
The preprint notes that real-world video benchmarks offer broad coverage, but their fixed clips mix together event count, rate, duration, and visual complexity, making it hard to pinpoint why a model fails. Existing programmatically generated benchmarks control these factors better, but they only score the final answer rather than checking reported events against executable ground truth. To close this gap, the authors introduce a trace-grounded parametric profiling benchmark that can systematically vary these factors and verify reported events against checkable ground truth. Experiments using this benchmark show that once confounding factors are properly separated, video language models struggle even with the basic task of accurately counting or tracking distinct events within a video. This finding suggests that reported performance on entangled real-world benchmarks may have masked a more fundamental weakness. The result points to event counting as a meaningful diagnostic for evaluating video understanding models.
Classification
Evidence 1
- The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping arXiv (cs.AI) 2026-08-06 accessed 2026-08-10T08:20:08+00:00
Part of trends 0
No objects.
Directly linked issues 0
No objects.
Public id: fm-bad84b7489da
