Sep 2026· Systems research and behavioral science· 0 citations· 24 references
TL;DR
Results show that model rankings depend on the evaluation protocol: claude‐haiku‐4.5 achieves the highest adjusted overall score in the human evaluation, whereas gpt‐5‐mini achieves the highest aggregate scores in the automated LLM‐as‐judge evaluation and the lowest estimated operational cost.
Abstract
This study proposes an AI‐assisted peer‐review framework based on retrieval‐augmented generation (RAG), advanced prompt engineering and multi‐agent large language model (LLM) orchestration to support human reviewers and editorial decision‐making through structured, evidence‐aware manuscript assessment. Manuscripts are represented as structured analytical objects connected to vectorized corpora of Q1/Q2 journal literature and reviewer guidelines. The framework employs five specialized reviewer agents: structural, methodological, incremental novelty, disruptive novelty and claim verification reviewers, using role‐conditioned and retrieval‐grounded prompts to produce schema‐constrained analytical outputs. These outputs are synthesized through a meta‐review before being delivered to a human reviewer. The system is evaluated on 100 manuscripts. AI‐assisted reviews are generated with
gpt‐5‐mini
,
claude‐haiku‐4.5
and
gemini‐2.5‐flash
, while review quality is assessed through blinded human evaluation and an automated LLM‐as‐judge procedure implemented with
grok‐4.3
. Results show that model rankings depend on the evaluation protocol:
claude‐haiku‐4.5
achieves the highest adjusted overall score in the human evaluation, whereas
gpt‐5‐mini
achieves the highest aggregate scores in the automated LLM‐as‐judge evaluation and the lowest estimated operational cost. Retrieval‐grounded prompts produced small but consistent gains across the six evaluated quality dimensions, including a 0.10 increase in
adjusted_overall_score
, while the no‐RAG variant remained highly competitive.
Automatic scientific survey generation has become an important task in scientific document processing. The common approach of retrieving literature from a single source (e.g., arXiv) and generating surveys through a one-pass large language model (LLM) call often leads to limited reference coverage and, more importantly...
Tong Bao, Mir Tafseer Nayeem, Yi Zhao et al.· Knowledge-Based Systems· 0 citations
Metag is a dataset to accelerate the development of meta-reviewing agents, specifically to identify changes made to scientific articles during the review-rebuttal process and will enable building methods to empower meta reviewers to quickly identify whether authors have addressed reviewer statements and where in the pa...
Anirudh S. Sundar, Min Chen, Divya Tadimeti et al.· 0 citations
A Task-Contingent Legitimacy framework offering a task-tiered policy approach and testable propositions is formalized in a Task-Contingent Legitimacy framework offering a task-tiered policy approach and testable propositions.
Eungi Kim, Vaishali Singh· Publications· 0 citations
Generating coherent meta-reviews from multiple peer reviews is challenging when reviewer evidence conflicts and varies in reliability. Existing approaches typically formulate meta-review generation as a multi-document summarization task and aggregate reviewer feedback uniformly, making it difficult to determine which o...
Xin-Zhe Wang, Fei Tao, Jiang Xie et al.· 0 citations
The growing volume of scientific output and the pace of AI advancement create a dual challenge: traditional systematic reviews take twelve to eighteen months to complete, risking obsolescence before publication, while the technology needed to accelerate them is itself advancing faster than it can be reviewed. This work...
E. A. Merchán-Cruz, Ioseb Gabelaia, Shwe Soe et al.· Informatics· 0 citations
CoSLR is presented, a Human-AI collaborative multi-agent system that supports the SLR workflow through a modular three-phase pipeline using large language models and Retrieval-Augmented Generation, and that places explicit, mandatory human checkpoints on the path between generated output and its acceptance.
Aidul Islam, M. Sami, Muhammad Waseem et al.· 0 citations