Skip to content
Open access

A Comparative Analysis of Large Language Model Performance on USMLE Step 1-Style Allergy/Immunology Questions: Evaluating Correctness and Consistency

Mar 2026 · Applied Clinical Informatics · Vol 17, pp. 754 - 762 · 0 citations · 55 references
Medicine

TL;DR

On Allergy/Immunology Step 1-style questions, Gemini and Grok demonstrated higher accuracy than ChatGPT, although their overall accuracies remained approximately 81%.

Abstract

Abstract Background Large language models (LLMs) are rapidly transforming medical education, yet their performance in Allergy/Immunology remains insufficiently characterized. Furthermore, concerns regarding accuracy, consistency, and sensitivity to input format persist. Objectives This study aimed to evaluate and compare the accuracy and response consistency of three leading LLMs—ChatGPT-5, Gemini 2.5, and Grok 4—on Allergy/Immunology United States Medical Licensing Examination (USMLE) Step 1-style questions under different prompt conditions. Methods Thirty-five USMLE Step 1-style questions were selected. Questions were presented to each model in two formats: single-question prompts and a combined prompt containing all questions. Fifteen trials were conducted for each format per model. Performance was assessed using mean accuracy, and variability was measured using Shannon entropy. Mixed-effects models tested the effects of model, prompt condition, and question difficulty. Results Overall accuracy differed significantly ( p < 0.001), with Gemini (80.7%) and Grok (80.5%) achieving higher mean scores than ChatGPT (74.3%). Single-item prompts yielded superior performance, with Grok (93.1%) and Gemini (90.9%) demonstrating the highest accuracy. Transitioning to a combined prompt significantly reduced accuracy for all models. Accuracy also decreased with increasing question difficulty for all models. Grok demonstrated superior reliability, maintaining the lowest overall response entropy, whereas ChatGPT exhibited the highest variability. Conclusion On Allergy/Immunology Step 1-style questions, Gemini and Grok demonstrated higher accuracy than ChatGPT, although their overall accuracies remained approximately 81%. Grok offered the most consistent performance. All models demonstrated substantial sensitivity to prompt complexity and inherent performance limitations. These findings underscore the importance of prompt optimization and support the supplementary role of these models in medical education.

Read PDF

Similar papers

Open access Sep 2026

Performance of large language models in answering public questions about nutrition in cirrhosis: a comparative study.

BACKGROUND large language models (LLMs) are increasingly used for public health information, but their performance in nutrition advice for cirrhosis remains uncertain. We investigated four LLMs in answering public questions about cirrhosis nutrition across safety, accuracy, empathy, information reliability and quality,...

Jun-Zheng Li, Ying-Jie Wu, Man Yang et al. · 0 citations
Review Open access Sep 2026

Benchmarking five large language models in medical genetics: a bilingual comparative evaluation using published and novel expert-authored questions

In an exploratory analysis, residual errors showed concentrated and concordant patterns, clustering in multistep Bayesian reasoning and evolving facts, suggesting that a second model may provide limited independent protection on such items.

Özge Beyza Gündoğdu Öğütlü, Benjamin D. Solomon, Y. Çelik · 0 citations
Review Open access Aug 2026

A Comparative Analysis of Large Language Models (LLMs) in Generating High-Quality Pharmacology Questions for Undergraduate Medical Education

The National Medical Commission embraced competency-based medical education (CBME) in 2019, which has a strong impact on higher-order cognitive abilities, professionalism, and clinical competence. Pharmacology in Phase II MBBS requires such competency-based evaluations. Large language models (LLMs) might help in...

Gurusamy Sivagnanam, Parasuraman Jayabharathy, Ganesan Justina Princess · 0 citations
Open access Oct 2026

Large language models’ performance in answering common patient questions about colonoscopy: an expert-based evaluation

This study aimed to compare the performance of two large language models (LLMs) the ChatGPT-4 and the Google Bard in delivering medical information about colonoscopy. It also sought to evaluate the models’ responses to common patient and caregiver questions in terms of accuracy, applicability, comprehensiveness, and co...

Neslihan Güneş Aydemi̇r, D. Yapar, Yasemin Demi̇r Avcı et al. · 0 citations
#software testing Review Open access Oct 2026

Performance of large language models in prescription review: a comparative study

Large language models (LLMs) have emerged as a novel technical solution for intelligent prescription review. Nevertheless, comparative analyses of their performance remain limited. This study comprehensively evaluates five mainstream LLMs (ChatGPT, DeepSeek, Kimi, Qwen and Doubao) in prescription review, aiming to...

Can Huang, Yan-Fang Sun, Wei Liu · 0 citations
Sep 2026

Beyond accuracy: an educational benchmarking study of task fragility and reasoning stability of large language models on dermatology board-style questions.

BACKGROUND Large language models (LLMs) are increasingly used in medical education and assessment frameworks. However, their reliability under varying task demands remains unclear. Significant gaps exist in the literature regarding the reasoning stability of these models when faced with fluctuations in question structu...

A. Atılan, Nıyazı Çetın · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.