Jul 2026· Health Informatics Journal· Vol 32 3, pp.
14604582261470637
· 0 citations· 21 references
Computer ScienceMedicine
TL;DR
AI models demonstrate generally high accuracy in medical examinations; however, they struggle with interpreting clinical context, recognising atypical medical scenarios and answering questions involving visual content.
Abstract
ObjectiveThis study evaluates the performance of large language models (LLMs)-ChatGPT-4.0, Gemini 2.0 Pro, o3-mini, Doctor GPT and DeepSeek-V3-in a national orthopaedic proficiency examination and explores their implications for health informatics and medical education. The responses of these models were analysed to assess accuracy rates and differences between models.MethodA total of 100 multiple-choice questions from the 2024 TOTEK examination were administered to each AI model under identical conditions. Correct and incorrect responses were recorded, and differences in performance were evaluated using chi-square testing and frequency analysis. Question categories were also compared to identify domain-specific variations.Resultso3-mini achieved the highest accuracy rate (79%), while Gemini 2.0 showed the lowest (68%); all models exceeded the 60% pass threshold. A statistically significant difference between models was identified in the Surgical Procedures category, in which Gemini 2.0 answered fewer questions correctly (23/36) than the other models (30-32/36) (χ2 = 9.87, df = 4, p = 0.043). No significant differences were observed in the remaining categories (all p > 0.05), and the overall difference in accuracy between models did not reach statistical significance (χ2 = 4.01, df = 4, p = 0.405). Clinical decision-making and visual content-based questions were the most challenging for all models.ConclusionAI models demonstrate generally high accuracy in medical examinations; however, they struggle with interpreting clinical context, recognising atypical medical scenarios and answering questions involving visual content.
AIM
Artificial intelligence (AI) has become an integral part of dental education and clinical practice. While several studies have assessed the performance of ChatGPT in medical exams, comparative analyses of various large language models (LLMs) in dentistry remain scarce. This study aimed to evaluate and compare the p...
N. Acar, Fatih Sengul, Periş Çelikel et al.· European journal of dental e...· 0 citations
LLMs demonstrate strong potential as supportive tools in dental education, particularly for text-based knowledge assessment, but their limited performance in visual question contexts highlights an important constraint for their integration into image-dependent domains such as dental diagnostics.
Sedef Ayşe Taşyapan, Didem Özer· European journal of dental e...· 0 citations
There may be significant differences in how effectively LLMs support patients with surgical queries, particularly in areas needing detailed explanation, and usually required minimal clarification in areas needing detailed explanation.
TP Davis, B. Guevel, K. Logishetty et al.· Annals of the Royal College...· 0 citations
Background The rapid development of large language models (LLMs) has raised questions about the continued educational value of conventional assessment formats in health professions education. This study evaluated how examination type, search access, image-based question characteristics, and item format influenced LLM a...
Toshitsugu Sakurai, Daichi Aizawa, Kazuyoshi Okawa et al.· Journal of Medical Education...· 0 citations
This study aimed to comprehensively evaluate the performance of seven leading large language models (LLMs) from the 2024-2025 period on Turkish Medical Specialty Examination (TUS) ophthalmology questions, analyzing factors such as question type, chronology, and clinical area, while critically assessing the risk of data...
Burhan Başkan, Yusuf Evcimen, Özgür Bülent Timuçin et al.· Muş Alparslan Üniversitesi S...· 0 citations
High-quality radiology assessment questions are essential for education competency evaluation but labor-intensive to create.
To compare four large language models (LLMs) in generating and evaluating radiology objective structured clinical examination (OSCE)–style questions and responses.
Fifty Radio...
Ankush Ankush, Samriddhi Burman, Sydney Smith et al.· Radiology Advances· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.