Signal New 'shadow evaluation' method finds AI agents can execute AI-research engineering but not open-ended research questions
Summary
Researchers introduced a new evaluation method called 'shadow evaluations' to measure progress toward AI research and development automation, addressing a gap where existing evaluations either test agents on narrow verifiable tasks or submit AI-generated papers to overstretched peer review. In this method, an AI agent takes on the central open-ended research question of a high-quality unpublished paper, and the paper's original authors grade the output. The team ran shadow evaluations on two unpublished NeurIPS 2026 submissions, giving frontier agents six days and thousands of dollars of compute. The agents completed all the engineering work without human help but could not make substantial progress toward answering the underlying research questions, and both papers were unambiguously rejected by their original authors. The researchers identified five recurring failure modes: poor judgment about the bar for publishable research, uncreative responses to research-design shortcomings, ineffective backtracking from dead ends, poor resource awareness, and instruction drift, with a robustness check using a second model and scaffold reproducing the same failures.
Classification
Evidence 1
- arXiv (cs.AI/cs.CY/cs.LG) 2026-07-29 accessed 2026-07-31T01:34:03+00:00
Part of trends 0
No objects.
Directly linked issues 0
No objects.
Public id: fm-1267212a3843