Skip to content
#small language model Review Open access

Who Reviews the Reviewer? A Multi‐Agent LLM Architecture With Meta‐Review Synthesis for Editorial Peer Review

Sep 2026 · Systems research and behavioral science · 0 citations · 24 references

TL;DR

Results show that model rankings depend on the evaluation protocol: claude‐haiku‐4.5 achieves the highest adjusted overall score in the human evaluation, whereas gpt‐5‐mini achieves the highest aggregate scores in the automated LLM‐as‐judge evaluation and the lowest estimated operational cost.

Abstract

This study proposes an AI‐assisted peer‐review framework based on retrieval‐augmented generation (RAG), advanced prompt engineering and multi‐agent large language model (LLM) orchestration to support human reviewers and editorial decision‐making through structured, evidence‐aware manuscript assessment. Manuscripts are represented as structured analytical objects connected to vectorized corpora of Q1/Q2 journal literature and reviewer guidelines. The framework employs five specialized reviewer agents: structural, methodological, incremental novelty, disruptive novelty and claim verification reviewers, using role‐conditioned and retrieval‐grounded prompts to produce schema‐constrained analytical outputs. These outputs are synthesized through a meta‐review before being delivered to a human reviewer. The system is evaluated on 100 manuscripts. AI‐assisted reviews are generated with gpt‐5‐mini , claude‐haiku‐4.5 and gemini‐2.5‐flash , while review quality is assessed through blinded human evaluation and an automated LLM‐as‐judge procedure implemented with grok‐4.3 . Results show that model rankings depend on the evaluation protocol: claude‐haiku‐4.5 achieves the highest adjusted overall score in the human evaluation, whereas gpt‐5‐mini achieves the highest aggregate scores in the automated LLM‐as‐judge evaluation and the lowest estimated operational cost. Retrieval‐grounded prompts produced small but consistent gains across the six evaluated quality dimensions, including a 0.10 increase in adjusted_overall_score , while the no‐RAG variant remained highly competitive.

Read PDF

Similar papers

SurveyAgent-HKA: A multi-agent framework for scientific survey generation with LLMs and human knowledge augmentation

Automatic scientific survey generation has become an important task in scientific document processing. The common approach of retrieving literature from a single source (e.g., arXiv) and generating surveys through a one-pass large language model (LLM) call often leads to limited reference coverage and, more importantly...

Tong Bao, Mir Tafseer Nayeem, Yi Zhao et al. · 0 citations
Review Aug 2026

Metag: A dataset to build agentic meta-reviewing capabilities

Metag is a dataset to accelerate the development of meta-reviewing agents, specifically to identify changes made to scientific articles during the review-rebuttal process and will enable building methods to empower meta reviewers to quickly identify whether authors have addressed reviewer statements and where in the pa...

Anirudh S. Sundar, Min Chen, Divya Tadimeti et al. · 0 citations
Review Open access Sep 2026

Artificial Intelligence in Peer Review: A Bibliometric-Guided Thematic Review and a Task-Contingent Legitimacy Framework

A Task-Contingent Legitimacy framework offering a task-tiered policy approach and testable propositions is formalized in a Task-Contingent Legitimacy framework offering a task-tiered policy approach and testable propositions.

Eungi Kim, Vaishali Singh · 0 citations

When Evidence Conflicts: Reliability-aware Meta-review Generation

Generating coherent meta-reviews from multiple peer reviews is challenging when reviewer evidence conflicts and varies in reliability. Existing approaches typically formulate meta-review generation as a multi-document summarization task and aggregate reviewer feedback uniformly, making it difficult to determine which o...

Xin-Zhe Wang, Fei Tao, Jiang Xie et al. · 0 citations
Review Open access Oct 2026

An Agentic AI-Aided Review of Large Language Model Applications in Literature Reviews: A Seven-Layer Architecture

The growing volume of scientific output and the pace of AI advancement create a dual challenge: traditional systematic reviews take twelve to eighteen months to complete, risking obsolescence before publication, while the technology needed to accelerate them is itself advancing faster than it can be reviewed. This work...

E. A. Merchán-Cruz, Ioseb Gabelaia, Shwe Soe et al. · 0 citations
#artificial intelligence Review Sep 2026

Checkpoints Are Not Enough: Trust Calibration in CoSLR, a Human-AI System for Systematic Literature Reviews

CoSLR is presented, a Human-AI collaborative multi-agent system that supports the SLR workflow through a modular three-phase pipeline using large language models and Retrieval-Augmented Generation, and that places explicit, mandatory human checkpoints on the path between generated output and its acceptance.

Aidul Islam, M. Sami, Muhammad Waseem et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.