Diagnostic capability of large language models in critically ill patients: a prospective single-centre study comparing ChatGPT, Claude, and Gemini with emergency physicians.
The LLMs achieved moderate agreement with ED reference diagnoses in critically ill patients but were consistently outperformed by physicians at the early diagnostic phases; their current diagnostic role in the ED remains limited.
Abstract
Background
Clinical decision-making requires integrating history, physical examination, laboratory, and imaging data. In the emergency department (ED), workload, time pressure, and cognitive burden may impair this process and affect decision quality. This study compares the diagnostic outputs of ChatGPT, Claude, and Gemini with those of emergency physicians in real-world ED cases.
Methods
This prospective, single-centre observational diagnostic agreement study compared the stage-wise outputs of four Large Language Models (LLMs) (ChatGPT-4o, ChatGPT-5, Claude Opus 4.1, and Gemini 2.5 Pro) with those of emergency physicians in critically ill ED patients. Between 10 August and 10 September 2025, de-identified clinical data were entered into the models via their official web interfaces using standardised prompts. In the first stage, physicians and LLMs each generated five preliminary diagnoses based on vital signs and medical history. In the second stage, following physical examination and laboratory and imaging results, both refined their lists into three differential diagnoses. In the third stage, the physicians' final diagnosis was accepted as the reference, and each LLM was prompted to provide a final diagnosis. LLM preliminary and differential diagnoses were compared with those of the physicians at the corresponding stage, and LLM final diagnoses with the reference; the inclusion of the final diagnosis within earlier lists was also evaluated. Agreement was quantified using Cohen's κ; analyses were performed in R.
Results
Of 389 screened patients, 180 were included (56.1% male; mean age 67 ± 15.9 years). Physicians contained the reference diagnosis within their top-5 preliminary and top-3 differential lists in 83.9% and 98.3% of cases, respectively, significantly exceeding every LLM (all p < 0.001). Final-diagnosis match rates were 67.2% [60.3-73.5] for ChatGPT-4o, 65.6% [58.7-71.9] for ChatGPT-5, 63.3% [56.3-69.9] for Claude Opus 4.1, and 59.4% [52.3-66.1] for Gemini 2.5 Pro (p = 0.16). Cohen's κ ranged from 0.575 (Gemini 2.5 Pro) to 0.656 (ChatGPT-4o), indicating moderate-to-substantial agreement, with no pairwise difference reaching significance.
Conclusions
The LLMs achieved moderate agreement with ED reference diagnoses in critically ill patients but were consistently outperformed by physicians at the early diagnostic phases. Despite final-diagnosis match rates of 59%-67%, their current diagnostic role in the ED remains limited.
Large language models outperformed emergency medicine physicians in overall diagnostic accuracy and demonstrated superior consistency across different types of diagnostic tasks, indicating that AI possesses robust pattern-recognition and reasoning capabilities, suggesting it could serve as a highly reliable clinical de...
Mehdi Arzani Shamsabadi, Roya Vatankhah, Hasan Jalilvand et al.· International Journal of Eme...· 0 citations
The hypMOOVE-PHC study is the hypothetical vignette phase of the Massive Open Online Validation and Evaluation (MOOVE) initiative, implemented in Kenya, Malawi, and Tanzania, and aims to validate a pool of LLMs through clinical review of expert-generated vignettes through fully crossed repeated-measures comparative eva...
P. Macharia, C. Kachimanga, M. Mahende et al.· medRxiv· 0 citations
The VLM demonstrated reliable clinical interpretation and an acceptable safety profile, however its integration into clinical workflows for early recognition of physiological deterioration and patient acuity assessment requires further rigorous evaluation and comparison to currently used track-and-trigger systems and p...
I. Strechen, P. Krishnan, O. Kilickaya et al.· International Journal of Med...· 0 citations
ChatGPT-5.2 had the highest observed performance across several predefined outcomes, although absolute differences were modest for some measures, particularly MCQ accuracy.
Although ChatGPT-5 demonstrated limited diagnostic performance and the ability to generate clinically interpretable explanations, its low specificity and limited agreement with a rule-based system highlight important safety concerns, these findings suggest that LLMs may serve as complementary tools rather than standalo...
Ilkay Ceylan, Serpil Ekin, Buket Özyaprak et al.· Frontiers in Pharmacology· 0 citations
RESPClinBench identifies task-specific limitations in multimodal pulmonary nodule assessment and longitudinal COPD management by combining explicit clinical-action coverage, holistic evaluation, and independent safety flags provides a clinically grounded basis for model selection and prospective validation.
Mou-Xiao Bian, Zhi Chen, Ruiyao Chen et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.