Using the ICLR 2025 review process, this study compares 2401 human reviews with 7203 reviews produced in separate, context-isolated API runs using Claude Sonnet 4.5, GPT-5.2 Thinking, and Gemini 3 Pro Preview across decision agreement, review-text characteristics, inter-model consistency, and human–AI aggregation.
Zhi-He Yang, Xiao-Yue Zhou, Hong-Sa Wang et al.· Publications· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.