Future Monitor 한국어

Signal Benchmark Finds AI Judges of Computer-Use Agents Systematically Mislabel Failures as Successes

Summary

A preprint posted July 30, 2026 introduces OSReward, the first systematic benchmark evaluating whether vision-language models (VLMs) can reliably judge whether computer-using AI agents (CUAs) completed their assigned tasks. As human verification cannot scale, the field increasingly relies on VLMs as judges of agent trajectories, but the paper finds this reliability had gone largely unexamined. Built from diverse agent trajectories across platforms with human-verified instructions and multi-stage ground-truth labeling, the benchmark yields two variants — OSReward-Hard, concentrating genuinely difficult cases, and OSReward-Multi, for fine-grained efficiency and alignment scoring. Across the most comprehensive evaluation of VLM judges to date, even state-of-the-art models fall short of an ideal judge and share a systematic leniency bias that mislabels failed agent runs as successes; the few models reliable enough to trust are too costly to run at scale, while affordable open models trail far behind. To close the gap, the authors release OS-Shepherd-100K, an open corpus of reasoning-annotated trajectory judgments, and train OS-Shepherd (9B and 35B parameter) open reward models that match commercial judges at 30-60% lower cost.

Classification

Main topicAI & Computing
Region menusGlobal
Impactscope:global
Time horizon0-3 years (2026-08-03)
Last updated2026-08-03T05:26:49.555560+00:00

Evidence 1

Part of trends 0

No objects.

Directly linked issues 0

No objects.

Public id: fm-0bdf909744bf