Mar 2026· Applied Clinical Informatics· Vol 17, pp. 754 - 762· 0 citations· 55 references
Medicine
TL;DR
On Allergy/Immunology Step 1-style questions, Gemini and Grok demonstrated higher accuracy than ChatGPT, although their overall accuracies remained approximately 81%.
Abstract
Abstract Background Large language models (LLMs) are rapidly transforming medical education, yet their performance in Allergy/Immunology remains insufficiently characterized. Furthermore, concerns regarding accuracy, consistency, and sensitivity to input format persist. Objectives This study aimed to evaluate and compare the accuracy and response consistency of three leading LLMs—ChatGPT-5, Gemini 2.5, and Grok 4—on Allergy/Immunology United States Medical Licensing Examination (USMLE) Step 1-style questions under different prompt conditions. Methods Thirty-five USMLE Step 1-style questions were selected. Questions were presented to each model in two formats: single-question prompts and a combined prompt containing all questions. Fifteen trials were conducted for each format per model. Performance was assessed using mean accuracy, and variability was measured using Shannon entropy. Mixed-effects models tested the effects of model, prompt condition, and question difficulty. Results Overall accuracy differed significantly ( p < 0.001), with Gemini (80.7%) and Grok (80.5%) achieving higher mean scores than ChatGPT (74.3%). Single-item prompts yielded superior performance, with Grok (93.1%) and Gemini (90.9%) demonstrating the highest accuracy. Transitioning to a combined prompt significantly reduced accuracy for all models. Accuracy also decreased with increasing question difficulty for all models. Grok demonstrated superior reliability, maintaining the lowest overall response entropy, whereas ChatGPT exhibited the highest variability. Conclusion On Allergy/Immunology Step 1-style questions, Gemini and Grok demonstrated higher accuracy than ChatGPT, although their overall accuracies remained approximately 81%. Grok offered the most consistent performance. All models demonstrated substantial sensitivity to prompt complexity and inherent performance limitations. These findings underscore the importance of prompt optimization and support the supplementary role of these models in medical education.
BACKGROUND
large language models (LLMs) are increasingly used for public health information, but their performance in nutrition advice for cirrhosis remains uncertain. We investigated four LLMs in answering public questions about cirrhosis nutrition across safety, accuracy, empathy, information reliability and quality,...
Jun-Zheng Li, Ying-Jie Wu, Man Yang et al.· Nutrición Hospitalaria· 0 citations
In an exploratory analysis, residual errors showed concentrated and concordant patterns, clustering in multistep Bayesian reasoning and evolving facts, suggesting that a second model may provide limited independent protection on such items.
Özge Beyza Gündoğdu Öğütlü, Benjamin D. Solomon, Y. Çelik· Frontiers in Medicine· 0 citations
The National Medical Commission embraced competency-based medical education (CBME) in 2019, which has a strong impact on higher-order cognitive abilities, professionalism, and clinical competence. Pharmacology in Phase II MBBS requires such competency-based evaluations. Large language models (LLMs) might help in...
Gurusamy Sivagnanam, Parasuraman Jayabharathy, Ganesan Justina Princess· The Journal of medical resea...· 0 citations
This study aimed to compare the performance of two large language models (LLMs) the ChatGPT-4 and the Google Bard in delivering medical information about colonoscopy. It also sought to evaluate the models’ responses to common patient and caregiver questions in terms of accuracy, applicability, comprehensiveness, and co...
Neslihan Güneş Aydemi̇r, D. Yapar, Yasemin Demi̇r Avcı et al.· BMC Gastroenterology· 0 citations
Large language models (LLMs) have emerged as a novel technical solution for intelligent prescription review. Nevertheless, comparative analyses of their performance remain limited. This study comprehensively evaluates five mainstream LLMs (ChatGPT, DeepSeek, Kimi, Qwen and Doubao) in prescription review, aiming to...
Can Huang, Yan-Fang Sun, Wei Liu· Frontiers in Medicine· 0 citations
BACKGROUND
Large language models (LLMs) are increasingly used in medical education and assessment frameworks. However, their reliability under varying task demands remains unclear. Significant gaps exist in the literature regarding the reasoning stability of these models when faced with fluctuations in question structu...
A. Atılan, Nıyazı Çetın· Cutaneous and Ocular Toxicol...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.