Results show that broad decision alignment does not imply agreement with finer human judgments or reviewing priorities, and cross-provider analysis of three complementary dimensions contributes a cross-provider analysis of three complementary dimensions.
Abstract
Large language models (LLMs) are increasingly used to generate scientific reviews, yet existing evaluations rarely examine whether different providers align with both conference decisions and human reviewing priorities within the same controlled setting. We compare reviews from OpenAI GPT-5.4, Google Gemini 3.1 Pro Preview, and Anthropic Claude Opus 4.6 with human reviews and final decisions for 300 topic-matched ICLR 2026 submissions, equally divided among oral, poster, and rejected papers. Each model reviewed every paper using identical instructions and rating scales after decision information was removed. Our study contributes a cross-provider analysis of three complementary dimensions: alignment with broad and fine-grained decision categories, differences in recommendation-scale usage, and thematic agreement in identified weaknesses. All three LLMs distinguished accepted from rejected papers, but none reproduced the oral versus poster distinction present in human ratings. Scoring patterns were provider-specific: Gemini assigned systematically higher ratings, while OpenAI and Claude were closer to humans for rejected and poster papers but more critical of oral papers. Human and LLM reviews also differed in emphasis, with LLMs more frequently identifying missing baseline comparisons and humans more often raising computational-efficiency concerns. These results show that broad decision alignment does not imply agreement with finer human judgments or reviewing priorities.
Using the ICLR 2025 review process, this study compares 2401 human reviews with 7203 reviews produced in separate, context-isolated API runs using Claude Sonnet 4.5, GPT-5.2 Thinking, and Gemini 3 Pro Preview across decision agreement, review-text characteristics, inter-model consistency, and human–AI aggregation.
Zhi-He Yang, Xiao-Yue Zhou, Hong-Sa Wang et al.· Publications· 0 citations
These findings show that LLM-assisted peer review changes the functional composition of review text, making it important to distinguish LLM-amplified critique from areas requiring human prioritization and accountable judgement.
YunHong Yang, Mike Thelwall, Guo-Xiu He· 0 citations
Peer-review feedback often arrives too late for authors to make meaningful revisions. We study an author-facing LLM system that moves part of this stress test before submission: it generates a broad pool of atomic concerns and compresses them into a short report. We evaluate agreement with historical reviews and, separ...
Findings show that many peer-review evaluation metrics partially conflate review quality with linguistic presentation, and indicate that robustness to meaning-preserving rewriting should be validated before such metrics are used to compare human-written, AI-assisted, and AI-generated reviews.
Shakiba Amirshahi, Sajad Ebrahimi, Hai-Son Le et al.· 0 citations
Grant application review is resource-intensive and subject to inter-rater variability. Large language models (LLMs) may augment this process, but their reliability in grant evaluation remains unexplored. This exploratory pilot study compared LLM-generated grant reviews to human expert reviews across three prompt engi...
Hants Williams, Jack Evan Lamberg, Eric M. Lamberg· Frontiers in Education· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.