Skip to content
Review Open access

Large Language Models in Peer Review: Decision Alignment, Review-Text Characteristics, and Human–AI Aggregation at ICLR 2025

Aug 2026 · Publications · 0 citations · 38 references

TL;DR

Using the ICLR 2025 review process, this study compares 2401 human reviews with 7203 reviews produced in separate, context-isolated API runs using Claude Sonnet 4.5, GPT-5.2 Thinking, and Gemini 3 Pro Preview across decision agreement, review-text characteristics, inter-model consistency, and human–AI aggregation.

Abstract

Large language models (LLMs) are increasingly employed in scholarly peer review, yet their suitability as autonomous evaluators remains uncertain. Using the ICLR 2025 review process, this study compares 2401 human reviews with 7203 reviews produced in separate, context-isolated API runs using Claude Sonnet 4.5, GPT-5.2 Thinking, and Gemini 3 Pro Preview across decision agreement, review-text characteristics, inter-model consistency, and human–AI aggregation. Raw LLM scores showed systematic leniency and score compression. A 0.1-point grid search identified thresholds of 6.2, 6.3, and 6.7 for Claude, GPT, and Gemini, respectively; repeated stratified cross-validation reproduced these thresholds. When applied without retuning to a stratified balanced sample of 300 ICLR 2024 papers, decision-agreement accuracy was 0.927, 0.913, and 0.930. Independent human coding of research type and primary field showed substantial pre-adjudication agreement (Cohen’s kappa = 0.774 and 0.714), and the recalculated analyses did not support H3. Review-text indicators showed similar structural completeness across sources but uneven critical-section length; these descriptive measures do not establish review quality. Human-containing aggregation rules showed higher agreement with conference decisions than corresponding AI-only rules, without establishing independent review quality or causal complementarity. A textual-overlap check found very low exact eight-gram containment, and manual inspection of the highest-similarity 1% found shared manuscript content or domain terminology rather than reviewer-specific evaluative language; possible prior exposure nevertheless could not be excluded.

Read PDF

Similar papers

Review Aug 2026

How Closely Do LLM Reviews Align with Human Peer Review?

Results show that broad decision alignment does not imply agreement with finer human judgments or reviewing priorities, and cross-provider analysis of three complementary dimensions contributes a cross-provider analysis of three complementary dimensions.

Abraham Camelo-Guerrero, J. Diaz-Rodriguez · 1 citation
Review Open access Aug 2026

Large language models as judges: recent advances in LLM-based evaluation, critique, preference modeling, and feedback for text and code

This survey provides a comprehensive overview of recent advances in LLM-based evaluation, covering techniques, applications, and challenges across domains, with future directions emphasizing standardized protocols, uncertainty estimation, and human–AI collaboration.

M. Nadăş · 0 citations
#artificial intelligence Review Jul 2026

Do large language models scrutinise what they review? A multimodal audit of scoring calibration, error detection, and author-identity effects

This study evaluates two multimodal LLMs, Qwen2.5-VL-72B and Pixtral-Large-124B, as reviewers across 165 submissions to the 2026 International Conference on Learning Representations, a venue that postdates both models'training cutoffs.

Emad Alharbi · 0 citations
Review Open access Aug 2026

LLM aspect prediction: reviewing academic papers from different aspects with Large Language Model

LLMAspectPrediction, a novel framework designed to predict fine-grained aspect scores for academic papers, is proposed, which assists reviewers by providing consistent, criteria-driven assessments and offers authors actionable feedback aligned with peer review standards.

Zi-Hao Hu, F. Fukumoto, Jian He et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.