Skip to content
Review Open access

Artificial intelligence in clinical decision-making: a comparison of ChatGPT 5.0 and Gemini 3.0 in otologic cases

Aug 2026 · BMC Medical Informatics and Decision Making · 0 citations

TL;DR

While both models show potential as clinical decision-support tools, Gemini 3.0 exhibited superior diagnostic performance in real-world otologic cases in the first study evaluating LLMs using real-world otologic data.

Abstract

Large Lingual Models (LLMs) can suggest treatment options based on a patient’s history and symptoms. However, they may lack the ability to deliver fully patient-specific recommendations, which may lead to misdiagnosis or inappropriate treatment decisions. This study compared the diagnostic accuracy, investigation recommendations, and treatment planning consistency of two LLMs, ChatGPT 5.0 and Gemini 3.0, using real-world otologic cases. Fifty retrospective cases with diverse clinical characteristics (symptoms, examination findings, audiometric tests, and radiological sections) were selected. Three experienced otorhinolaryngologists established the final diagnoses as a gold standard. The cases were processed by both LLMs, which were instructed to act as ENT specialists. Two blinded reviewers independently evaluated the responses using the Artificial Intelligence Performance Instrument (AIPI) across four domains: summarization, differential diagnosis, additional examination, and therapy options. Inter-rater reliability was assessed via Cohen’s kappa. Both models demonstrated high performance in managing otologic cases. However, Gemini 3.0 significantly outperformed ChatGPT 5.0 in diagnostic accuracy, investigation/therapy planning, and overall AIPI scores ( p  < 0.05). Subgroup analysis revealed that Gemini 3.0 performed higher AIPI scores in ‘vertigo’ cases, while similar performances were found in ‘otits media’, ‘hearing loss’ and ‘facial paralysis’ groups. Inter-rater agreement for AIPI scores was excellent. To our knowledge, this is the first study evaluating LLMs using real-world otologic data. While both models show potential as clinical decision-support tools, Gemini 3.0 exhibited superior diagnostic performance in real-world otologic cases. Future research should include larger populations and endoscopic imaging to further validate these findings.

Read PDF

Similar papers

Open access Aug 2026

Patient education for neuromyelitis optica spectrum disorder using large language models: combining expert assessment and real-world patient interaction

Objectives This study aimed to compare the educational performance of six mainstream LLMs for neuromyelitis optica spectrum disorder (NMOSD) and evaluated patient satisfaction during real-world interactions. Methods This study was conducted from March to April 2026. In the first Phase, Twenty NMOSD-related questions de...

Chen Li, Yu-Ting Hu, Xiao-Yan Wang et al. · 0 citations
Preprint Aug 2026

Improving O-RADS Risk Stratification from Ultrasound Reports: A Comparative Evaluation of Hybrid versus End-to-End LLM Reasoning Strategies

The feature-based hybrid LLM architecture that separates clinical feature extraction from deterministic guideline execution enables highly accurate, reliable, and interpretable automated O-RADS classification, providing a promising approach for standardized, guideline-based clinical decision-making.

Xiao-Tong Tan, Chunli Qiu, Xin Liu et al. · 0 citations
Aug 2026

Comparative evaluation of large language models and clinicians in real-world glaucoma clinical reasoning.

Large language model-based AI systems produced structured glaucoma-related reasoning with performance that overlapped with attending ophthalmologists but did not establish clinical equivalence, but may have potential as supervised decision-support and educational tools.

Hou-Fa Yin, Lixia Shen, Haiyan Cai et al. · 0 citations

Beyond diagnostic accuracy. Applying and extending methods for diagnostic test research

To assess the information content of both MRJ and CT imaging modalities using the entropy statistic, the information content for each test was calculated in terms of entropy using a random effects logistic regression model.

A. Glas, Nina M. Klemetsö, J. V. van Rijn et al. · 0 citations
Open access Sep 2026

Comparative Expert Evaluation of Multimodal Large Language Models for Pediatric Rash Diagnosis: Clinical Utility, Safety, Information Quality, and Readability

Background/Objectives: Multimodal large language models (LLMs) can interpret clinical text and images, but their performance in pediatric rash assessment remains uncertain. This study compared the clinical utility, safety, information quality, diagnostic correctness, and readability of ChatGPT, Gemini and Grok. Methods...

D. Lahut, Özlem Erdede, R. S. Sezer Yamanel · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.