Skip to content

Author

Mustafa Durmaz

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Review Sep 2026

Large Language Models as Peer Reviewers: Prompt Sensitivity and Model-Dependent Reproducibility.

RATIONALE AND OBJECTIVES To evaluate the reproducibility of editorial re--ations by Large Language Models, agreement across different models, prompt-sensitivity, and fidelity of critique statements to source manuscripts in a simulated peer-review setting. MATERIALS AND METHODS Fifteen open-access radiology manuscripts were anonymized and reviewed by eight large language models (LLMs) across four developer families (ChatGPT, DeepSeek, Gemini, Grok) with two different prompts. Each manuscript-model-prompt condition was repeated across three independent runs, yielding 720 reviews. Intra-model stability was defined as identical decisions across runs. Inter-model agreement was assessed with Fleiss' kappa. A stratified random sample of 128 reviews underwent manual verification against the source manuscripts and was categorized as grounded, distorted, or hallucinated. RESULTS Across all reviews, decisions were Minor Revision in 51.3% (369 of 720), Major Revision in 43.9% (316 of 720), Accept in 4.9% (35 of 720), and Reject in 0% (0 of 720). Prompt strictness shifted decision severity (p < 0.001): Prompt 1 yielded 9.7% Accept, 64.7% Minor Revision, and 25.6% Major Revision, whereas Prompt 2 eliminated Accept and increased Major Revision to 62.2%. Inter-model agreement was fair (Fleiss' κ = 0.25). In the audit, 94.0% of statements were grounded and 6.0% were distorted, with no hallucinated statements observed. CONCLUSION Large language model editorial re--ations were prompt sensitive and showed fair agreement across models despite critique statements that were largely grounded in manuscript text, supporting assistive use with human oversight.

S. Erturk, Mustafa Durmaz · 1 citation