Skip to content
Review

AI-Assisted Peer Review Across Research Communities: From Reviewer AI Policies to LLM Review Quality

Aug 2026 · 1 citation · 28 references
Computer Science

TL;DR

It is demonstrated that aggregate quality scores alone can overestimate review quality and argued for multi-dimensional evaluation of AI-generated peer reviews.

Abstract

AI-assisted peer review is increasingly discussed and adopted as a tool to support the scientific publishing process, yet there is little systematic understanding of how publication venues regulate its use or of how capable current AI review systems are. We address these questions by first surveying reviewer-facing AI policies across 111 leading AI/NLP conferences and medical journals, revealing substantial regulation differences between the two communities. Second, we evaluate AI-generated peer reviews at ICLR 2026 and Nature Communications using a novel dataset comprising original manuscript submissions and several hundred human- and machine-generated reviews. We compare reviews produced by open-source and proprietary models using complementary evaluation metrics, including LLM-as-a-Judge, score alignment, granularity, and overlap with human reviewers'concerns. Our results show that current LLMs can generate detailed and fluent reviews but exhibit systematic weaknesses, such as overly positive recommendations, generic criticism, and uneven evidence grounding. We demonstrate that aggregate quality scores alone can overestimate review quality and argue for multi-dimensional evaluation of AI-generated peer reviews.

View source

Similar papers

Review

Utility and Trustworthiness of Generative AI in Peer Review

The findings suggest that modern Large Language Models can provide useful and consistent support for scientific peer review, however remaining differences between AI-generated and human-generated evaluations indicate that current systems should be viewed as complementary tools that assist human reviewers rather than re...

Vuk D. Tomić, T. Heyman, E. V. van Nieuwenburg · 0 citations
#machine learning Review Sep 2026

When AI Reviews Train AI Reviewers: Scientific-Judgment Collapse and Mitigation

This study introduces TrustReviewer, an open-source LLM-based system for generating peer reviews of AI and machine learning papers and describes a concrete risk of recursive reviewer training and provides practical interventions for preserving judgment diversity and improving recommendation alignment in AI-assisted sci...

Sy-Tuyen Ho, Ming-Hui Liu, Fu-Rong Huang · 0 citations
Review Aug 2026

How Closely Do LLM Reviews Align with Human Peer Review?

Results show that broad decision alignment does not imply agreement with finer human judgments or reviewing priorities, and cross-provider analysis of three complementary dimensions contributes a cross-provider analysis of three complementary dimensions.

Abraham Camelo-Guerrero, J. Diaz-Rodriguez · 1 citation

Judging a Review by its Cover: A Reliability Analysis of LLM-based Peer Review Evaluation Metrics

Findings show that many peer-review evaluation metrics partially conflate review quality with linguistic presentation, and indicate that robustness to meaning-preserving rewriting should be validated before such metrics are used to compare human-written, AI-assisted, and AI-generated reviews.

Shakiba Amirshahi, Sajad Ebrahimi, Hai-Son Le et al. · 0 citations
Review Open access Aug 2026

Large Language Models in Peer Review: Decision Alignment, Review-Text Characteristics, and Human–AI Aggregation at ICLR 2025

Using the ICLR 2025 review process, this study compares 2401 human reviews with 7203 reviews produced in separate, context-isolated API runs using Claude Sonnet 4.5, GPT-5.2 Thinking, and Gemini 3 Pro Preview across decision agreement, review-text characteristics, inter-model consistency, and human–AI aggregation.

Zhi-He Yang, Xiao-Yue Zhou, Hong-Sa Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.