Signal AI-Assisted Peer Review Across Research Communities: From Reviewer AI Policies to LLM Review Quality
Summary
Alexander M. Fichtl, Lukas Ellinger, Josefin Kelber, Kryštof Olík, and Georg Groh published a paper on arXiv (cs.CY) on August 4, 2026, examining AI-assisted peer review across research communities. The authors first surveyed reviewer-facing AI policies across 111 leading AI/NLP conferences and medical journals, finding substantial regulatory differences between the two communities. Second, they evaluated AI-generated peer reviews at ICLR 2026 and Nature Communications using a novel dataset combining original manuscript submissions with several hundred human- and machine-generated reviews, comparing open-source and proprietary models with complementary metrics including LLM-as-a-Judge, score alignment, granularity, and overlap with human reviewers' concerns. The results show current LLMs can generate detailed, fluent reviews but exhibit systematic weaknesses such as overly positive recommendations, generic criticism, and uneven evidence grounding. The authors demonstrate that aggregate quality scores alone can overestimate review quality and argue for multi-dimensional evaluation of AI-generated peer reviews.
Classification
Evidence 1
- arXiv (cs.CY) 2026-08-04 accessed 2026-08-05T02:34:16+00:00
Part of trends 0
No objects.
Directly linked issues 0
No objects.
Public id: fm-7ea2af23e298