Skip to content
Open access

Accuracy, Reliability, and Bloom’s Taxonomy Performance of Seven Large Language Models on Microbiology Questions

Aug 2026 · Advances in Medical Education and Practice · Vol 17 · 0 citations · 19 references
Medicine

Abstract

Background Large language models (LLMs) are increasingly used as learning resources in medical education, yet their performance and reliability in microbiology, a discipline with a broad, heterogeneous knowledge base, have not been systematically evaluated across multiple platforms. Objective To benchmark seven publicly available LLMs on microbiology multiple-choice questions (MCQs), assessing overall accuracy, test-retest reliability, topic-specific performance, and the relationship between cognitive complexity and model performance. Methods Seven LLMs (Claude 4.6 Sonnet, Gemini 3.0, ChatGPT-5.2, Grok 4, Copilot, DeepSeek V3, and Kimi K2) completed 200 MCQs distributed across 20 microbiology topics and five Bloom’s taxonomy levels in three independent sessions separated by 24-hour intervals. A total of 4200 responses were analyzed. Statistical analysis included one-way ANOVA with Tukey’s HSD post hoc tests, repeated-measures ANOVA, intraclass correlation coefficients (ICCs), and Pearson correlations. Results The collective mean accuracy was 86.18%. Six of seven systems exceeded the 80% high-competency threshold; Claude (89.83%), Grok (89.50%), and GPT (88.83%) led the group. Gemini (72.33%) was the only underperforming system. Test-retest reliability varied dramatically: Claude achieved excellent ICC (0.966), while Gemini exhibited poor reliability (ICC = 0.290), with session-to-session fluctuations of up to 100 percentage points on individual topics. Microbial Cell (100%) was the easiest topic; Viral Genomics (61.9%) was the most challenging across all systems. A uniform decline at Bloom’s Level 4 (Analyze) was observed across all LLMs, with no model exceeding 78%. Conclusion Contemporary LLMs demonstrate substantial knowledge of microbiology but differ markedly in reliability. Response consistency, alongside accuracy, should be a primary criterion for educational deployment. These findings are specific to microbiology MCQ performance and may not generalize to open-ended clinical reasoning.

Read PDF