Skip to content
Review

Large Language Models as Peer Reviewers: Prompt Sensitivity and Model-Dependent Reproducibility.

Sep 2026 · Academic Radiology · Vol 33 9, pp. 3642-3650 · 1 citation · 36 references
Medicine

Abstract

Rationale

AND

Objectives

To evaluate the reproducibility of editorial re--ations by Large Language Models, agreement across different models, prompt-sensitivity, and fidelity of critique statements to source manuscripts in a simulated peer-review setting.

Materials And Methods

Fifteen open-access radiology manuscripts were anonymized and reviewed by eight large language models (LLMs) across four developer families (ChatGPT, DeepSeek, Gemini, Grok) with two different prompts. Each manuscript-model-prompt condition was repeated across three independent runs, yielding 720 reviews. Intra-model stability was defined as identical decisions across runs. Inter-model agreement was assessed with Fleiss' kappa. A stratified random sample of 128 reviews underwent manual verification against the source manuscripts and was categorized as grounded, distorted, or hallucinated.

Results

Across all reviews, decisions were Minor Revision in 51.3% (369 of 720), Major Revision in 43.9% (316 of 720), Accept in 4.9% (35 of 720), and Reject in 0% (0 of 720). Prompt strictness shifted decision severity (p < 0.001): Prompt 1 yielded 9.7% Accept, 64.7% Minor Revision, and 25.6% Major Revision, whereas Prompt 2 eliminated Accept and increased Major Revision to 62.2%. Inter-model agreement was fair (Fleiss' κ = 0.25). In the audit, 94.0% of statements were grounded and 6.0% were distorted, with no hallucinated statements observed.

Conclusion

Large language model editorial re--ations were prompt sensitive and showed fair agreement across models despite critique statements that were largely grounded in manuscript text, supporting assistive use with human oversight.

View source