Aug 2026· BMC Medical Informatics and Decision Making· 0 citations
TL;DR
While both models show potential as clinical decision-support tools, Gemini 3.0 exhibited superior diagnostic performance in real-world otologic cases in the first study evaluating LLMs using real-world otologic data.
Abstract
Large Lingual Models (LLMs) can suggest treatment options based on a patient’s history and symptoms. However, they may lack the ability to deliver fully patient-specific recommendations, which may lead to misdiagnosis or inappropriate treatment decisions.
This study compared the diagnostic accuracy, investigation recommendations, and treatment planning consistency of two LLMs, ChatGPT 5.0 and Gemini 3.0, using real-world otologic cases.
Fifty retrospective cases with diverse clinical characteristics (symptoms, examination findings, audiometric tests, and radiological sections) were selected. Three experienced otorhinolaryngologists established the final diagnoses as a gold standard. The cases were processed by both LLMs, which were instructed to act as ENT specialists. Two blinded reviewers independently evaluated the responses using the Artificial Intelligence Performance Instrument (AIPI) across four domains: summarization, differential diagnosis, additional examination, and therapy options. Inter-rater reliability was assessed via Cohen’s kappa.
Both models demonstrated high performance in managing otologic cases. However, Gemini 3.0 significantly outperformed ChatGPT 5.0 in diagnostic accuracy, investigation/therapy planning, and overall AIPI scores (
p
< 0.05). Subgroup analysis revealed that Gemini 3.0 performed higher AIPI scores in ‘vertigo’ cases, while similar performances were found in ‘otits media’, ‘hearing loss’ and ‘facial paralysis’ groups. Inter-rater agreement for AIPI scores was excellent.
To our knowledge, this is the first study evaluating LLMs using real-world otologic data. While both models show potential as clinical decision-support tools, Gemini 3.0 exhibited superior diagnostic performance in real-world otologic cases. Future research should include larger populations and endoscopic imaging to further validate these findings.
Objectives This study aimed to compare the educational performance of six mainstream LLMs for neuromyelitis optica spectrum disorder (NMOSD) and evaluated patient satisfaction during real-world interactions. Methods This study was conducted from March to April 2026. In the first Phase, Twenty NMOSD-related questions de...
Chen Li, Yu-Ting Hu, Xiao-Yan Wang et al.· Frontiers in Physiology· 0 citations
Large language model-based AI systems produced structured glaucoma-related reasoning with performance that overlapped with attending ophthalmologists but did not establish clinical equivalence, but may have potential as supervised decision-support and educational tools.
Hou-Fa Yin, Lixia Shen, Haiyan Cai et al.· Graefe's archive for clinica...· 0 citations
To assess the information content of both MRJ and CT imaging modalities using the entropy statistic, the information content for each test was calculated in terms of entropy using a random effects logistic regression model.
A. Glas, Nina M. Klemetsö, J. V. van Rijn et al.· 0 citations
Background/Objectives: Multimodal large language models (LLMs) can interpret clinical text and images, but their performance in pediatric rash assessment remains uncertain. This study compared the clinical utility, safety, information quality, diagnostic correctness, and readability of ChatGPT, Gemini and Grok. Methods...
D. Lahut, Özlem Erdede, R. S. Sezer Yamanel· Children· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.