Skip to content
Open access

Diagnostic capability of large language models in critically ill patients: a prospective single-centre study comparing ChatGPT, Claude, and Gemini with emergency physicians.

Jul 2026 · BMC Emergency Medicine · 0 citations
Medicine

TL;DR

The LLMs achieved moderate agreement with ED reference diagnoses in critically ill patients but were consistently outperformed by physicians at the early diagnostic phases; their current diagnostic role in the ED remains limited.

Abstract

Background

Clinical decision-making requires integrating history, physical examination, laboratory, and imaging data. In the emergency department (ED), workload, time pressure, and cognitive burden may impair this process and affect decision quality. This study compares the diagnostic outputs of ChatGPT, Claude, and Gemini with those of emergency physicians in real-world ED cases.

Methods

This prospective, single-centre observational diagnostic agreement study compared the stage-wise outputs of four Large Language Models (LLMs) (ChatGPT-4o, ChatGPT-5, Claude Opus 4.1, and Gemini 2.5 Pro) with those of emergency physicians in critically ill ED patients. Between 10 August and 10 September 2025, de-identified clinical data were entered into the models via their official web interfaces using standardised prompts. In the first stage, physicians and LLMs each generated five preliminary diagnoses based on vital signs and medical history. In the second stage, following physical examination and laboratory and imaging results, both refined their lists into three differential diagnoses. In the third stage, the physicians' final diagnosis was accepted as the reference, and each LLM was prompted to provide a final diagnosis. LLM preliminary and differential diagnoses were compared with those of the physicians at the corresponding stage, and LLM final diagnoses with the reference; the inclusion of the final diagnosis within earlier lists was also evaluated. Agreement was quantified using Cohen's κ; analyses were performed in R.

Results

Of 389 screened patients, 180 were included (56.1% male; mean age 67 ± 15.9 years). Physicians contained the reference diagnosis within their top-5 preliminary and top-3 differential lists in 83.9% and 98.3% of cases, respectively, significantly exceeding every LLM (all p < 0.001). Final-diagnosis match rates were 67.2% [60.3-73.5] for ChatGPT-4o, 65.6% [58.7-71.9] for ChatGPT-5, 63.3% [56.3-69.9] for Claude Opus 4.1, and 59.4% [52.3-66.1] for Gemini 2.5 Pro (p = 0.16). Cohen's κ ranged from 0.575 (Gemini 2.5 Pro) to 0.656 (ChatGPT-4o), indicating moderate-to-substantial agreement, with no pairwise difference reaching significance.

Conclusions

The LLMs achieved moderate agreement with ED reference diagnoses in critically ill patients but were consistently outperformed by physicians at the early diagnostic phases. Despite final-diagnosis match rates of 59%-67%, their current diagnostic role in the ED remains limited.

Read PDF

Similar papers

Open access Jul 2026

Comparing the text-based diagnostic reasoning performance of emergency medicine physicians and large language models in both definitive and differential diagnoses using standardized clinical vignettes: a preliminary study

Large language models outperformed emergency medicine physicians in overall diagnostic accuracy and demonstrated superior consistency across different types of diagnostic tasks, indicating that AI possesses robust pattern-recognition and reasoning capabilities, suggesting it could serve as a highly reliable clinical de...

Mehdi Arzani Shamsabadi, Roya Vatankhah, Hasan Jalilvand et al. · 0 citations
Review Open access Sep 2026

Evaluating large language models as clinical decision support tools in primary healthcare settings: Protocol for a multi-country comparative validation study on expert-adjudicated hypothetical vignettes (hypMOOVE-PHC)

The hypMOOVE-PHC study is the hypothetical vignette phase of the Massive Open Online Validation and Evaluation (MOOVE) initiative, implemented in Kenya, Malawi, and Tanzania, and aims to validate a pool of LLMs through clinical review of expert-generated vignettes through fully crossed repeated-measures comparative eva...

P. Macharia, C. Kachimanga, M. Mahende et al. · 0 citations
Review Aug 2026

Clinical evaluation of a vision-language model for optimizing triage and clinical workflows in critical care.

The VLM demonstrated reliable clinical interpretation and an acceptable safety profile, however its integration into clinical workflows for early recognition of physiological deterioration and patient acuity assessment requires further rigorous evaluation and comparison to currently used track-and-trigger systems and p...

I. Strechen, P. Krishnan, O. Kilickaya et al. · 0 citations
Open access Sep 2026

Laboratory Medicine Decision Support—Beyond Exam Passing: A Blinded 100-Case Text-Based Benchmark of Diagnostic Accuracy, Management Quality, and Safety for ChatGPT, Gemini, and DeepSeek—LLM Decision Support in Laboratory Medicine

ChatGPT-5.2 had the highest observed performance across several predefined outcomes, although absolute differences were modest for some measures, particularly MCQ accuracy.

K. Ulutaş, A. Pekmezci · 0 citations
Open access Sep 2026

Evaluating a large language model (ChatGPT-5) for detecting potential drug–drug interactions in intensive care: a cross-sectional comparative study with a clinical decision support system

Although ChatGPT-5 demonstrated limited diagnostic performance and the ability to generate clinically interpretable explanations, its low specificity and limited agreement with a rule-based system highlight important safety concerns, these findings suggest that LLMs may serve as complementary tools rather than standalo...

Ilkay Ceylan, Serpil Ekin, Buket Özyaprak et al. · 0 citations
Review Aug 2026

RESPClinBench: Benchmarking Multimodal Clinical Decision-Making and Longitudinal Disease Management in Respiratory Specialty Care

RESPClinBench identifies task-specific limitations in multimodal pulmonary nodule assessment and longitudinal COPD management by combining explicit clinical-action coverage, holistic evaluation, and independent safety flags provides a clinically grounded basis for model selection and prospective validation.

Mou-Xiao Bian, Zhi Chen, Ruiyao Chen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.