Skip to content
Review Open access

Robustness Gap of Large Language Models in Nephrology

Aug 2026 · medRxiv · 0 citations
Medicine

TL;DR

Newer models may be more robust, but multiple-choice accuracy remains an incomplete measure of clinical reasoning robustness, as all evaluated LLMs showed a significant robustness gap after NOTA replacement.

Abstract

Background: Whether benchmark performance reflects robust clinical reasoning rather than surface-level pattern recognition remains uncertain. We evaluated the robustness of state-of-the-art large language models (LLMs) on nephrology board renewal questions using "None of the other answers" (NOTA) substitution. Methods: From 210 Japanese Society of Nephrology board renewal questions (2014-2023), two nephrologists independently reviewed all items. Questions in which NOTA became the sole correct answer after replacement were included, yielding 145 validated questions. GPT-5, GPT-4o, Gemini 2.5 Pro, and Gemini 2.0 Flash were evaluated via application programming interfaces under default settings. The primary endpoint was accuracy, and paired differences were assessed using the exact two-sided McNemar test. Results: Accuracy was significantly lower after NOTA substitution for all models: GPT-4o, 66.21% to 19.31% (drop, 46.90 percentage points [pp]); GPT-5, 87.59% to 73.10% (14.48 pp); Gemini 2.0 Flash, 58.62% to 31.03% (27.59 pp); and Gemini 2.5 Pro, 86.90% to 55.86% (31.03 pp); all P < .001. GPT-5 showed the smallest decline and the highest accuracy in both versions. Conclusions: All evaluated LLMs showed a significant robustness gap after NOTA replacement. Newer models may be more robust, but multiple-choice accuracy remains an incomplete measure of clinical reasoning robustness.

Read PDF

Similar papers

Open access Jul 2026

Large language models for interpretation of health checkup results

Overall, advanced prompting markedly improved model performance, and top-tier LLMs demonstrated robust interpretive capability, while caution is needed for variables with complex clinical semantics such as blood pressure.

Jiwon You, Hangsik Shin · 0 citations
Review Aug 2026

Corrected: Safety in the Age of Artificial Intelligence: Evaluating Large Language Model Adherence to Antithrombotic Medication and Regional Anesthesia Guidelines

This study evaluates the accuracy of two of the foremost large language models (LLMs) used in clinical decision support tools for regional and neuraxial anesthesia procedures for patients receiving antithrombotic medications.

Conner M. Willson, Birpartap S. Thind, Jay Srinivas et al. · 0 citations
Open access Jul 2026

Evaluation of the performance and temporal variability of large language models in patient education regarding pneumothorax: a seven-day analysis.

While unprompted models exhibit marked baseline linguistic and quality variations, the strategic integration of robust prompt engineering successfully enforces the temporal stability and clarity required for reliable digital public health communication.

Ömer Önal, Suzan Temiz Bekce · 0 citations
Open access Jul 2026

Evaluating Large Language Models Against Clinical Assessment Frameworks for Early Sepsis Detection in the ICU

Timely recognition of sepsis remains difficult when early physiological abnormalities are subtle or incomplete. This study examined whether general-purpose large language models could discriminate sepsis risk from an initial ICU vital-sign snapshot as effectively as established clinical scoring approaches. We performed...

Anvit More, Vishala Bodetti, Kishan Gor et al. · 0 citations
Open access Sep 2026

Benchmarking ten frontier large language models on 1,477 board style multiple choice questions in hematology

Large Language Models (LLMs) are increasingly used by clinicians and patients for medical queries, yet their accuracy and safety at the specialist level in hematology remain insufficiently characterised. We benchmarked ten frontier proprietary and open-weight LLMs across two generations on 1,477 board-style hematology...

M. Radoynova, M. Benouis, F. Schulze et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.