It is demonstrated that aggregate quality scores alone can overestimate review quality and argued for multi-dimensional evaluation of AI-generated peer reviews.
Abstract
AI-assisted peer review is increasingly discussed and adopted as a tool to support the scientific publishing process, yet there is little systematic understanding of how publication venues regulate its use or of how capable current AI review systems are. We address these questions by first surveying reviewer-facing AI policies across 111 leading AI/NLP conferences and medical journals, revealing substantial regulation differences between the two communities. Second, we evaluate AI-generated peer reviews at ICLR 2026 and Nature Communications using a novel dataset comprising original manuscript submissions and several hundred human- and machine-generated reviews. We compare reviews produced by open-source and proprietary models using complementary evaluation metrics, including LLM-as-a-Judge, score alignment, granularity, and overlap with human reviewers'concerns. Our results show that current LLMs can generate detailed and fluent reviews but exhibit systematic weaknesses, such as overly positive recommendations, generic criticism, and uneven evidence grounding. We demonstrate that aggregate quality scores alone can overestimate review quality and argue for multi-dimensional evaluation of AI-generated peer reviews.
The findings suggest that modern Large Language Models can provide useful and consistent support for scientific peer review, however remaining differences between AI-generated and human-generated evaluations indicate that current systems should be viewed as complementary tools that assist human reviewers rather than re...
Vuk D. Tomić, T. Heyman, E. V. van Nieuwenburg· 0 citations
This study introduces TrustReviewer, an open-source LLM-based system for generating peer reviews of AI and machine learning papers and describes a concrete risk of recursive reviewer training and provides practical interventions for preserving judgment diversity and improving recommendation alignment in AI-assisted sci...
Results show that broad decision alignment does not imply agreement with finer human judgments or reviewing priorities, and cross-provider analysis of three complementary dimensions contributes a cross-provider analysis of three complementary dimensions.
Abraham Camelo-Guerrero, J. Diaz-Rodriguez· 1 citation
Findings show that many peer-review evaluation metrics partially conflate review quality with linguistic presentation, and indicate that robustness to meaning-preserving rewriting should be validated before such metrics are used to compare human-written, AI-assisted, and AI-generated reviews.
Shakiba Amirshahi, Sajad Ebrahimi, Hai-Son Le et al.· 0 citations
Using the ICLR 2025 review process, this study compares 2401 human reviews with 7203 reviews produced in separate, context-isolated API runs using Claude Sonnet 4.5, GPT-5.2 Thinking, and Gemini 3 Pro Preview across decision agreement, review-text characteristics, inter-model consistency, and human–AI aggregation.
Zhi-He Yang, Xiao-Yue Zhou, Hong-Sa Wang et al.· Publications· 0 citations
These findings show that LLM-assisted peer review changes the functional composition of review text, making it important to distinguish LLM-amplified critique from areas requiring human prioritization and accountable judgement.
YunHong Yang, Mike Thelwall, Guo-Xiu He· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.