Skip to content
Open access

Evaluation of large language models in a national orthopaedic proficiency examination: Implications for health informatics and medical education

Jul 2026 · Health Informatics Journal · Vol 32 3, pp. 14604582261470637 · 0 citations · 21 references
Computer Science Medicine

TL;DR

AI models demonstrate generally high accuracy in medical examinations; however, they struggle with interpreting clinical context, recognising atypical medical scenarios and answering questions involving visual content.

Abstract

ObjectiveThis study evaluates the performance of large language models (LLMs)-ChatGPT-4.0, Gemini 2.0 Pro, o3-mini, Doctor GPT and DeepSeek-V3-in a national orthopaedic proficiency examination and explores their implications for health informatics and medical education. The responses of these models were analysed to assess accuracy rates and differences between models.MethodA total of 100 multiple-choice questions from the 2024 TOTEK examination were administered to each AI model under identical conditions. Correct and incorrect responses were recorded, and differences in performance were evaluated using chi-square testing and frequency analysis. Question categories were also compared to identify domain-specific variations.Resultso3-mini achieved the highest accuracy rate (79%), while Gemini 2.0 showed the lowest (68%); all models exceeded the 60% pass threshold. A statistically significant difference between models was identified in the Surgical Procedures category, in which Gemini 2.0 answered fewer questions correctly (23/36) than the other models (30-32/36) (χ2 = 9.87, df = 4, p = 0.043). No significant differences were observed in the remaining categories (all p > 0.05), and the overall difference in accuracy between models did not reach statistical significance (χ2 = 4.01, df = 4, p = 0.405). Clinical decision-making and visual content-based questions were the most challenging for all models.ConclusionAI models demonstrate generally high accuracy in medical examinations; however, they struggle with interpreting clinical context, recognising atypical medical scenarios and answering questions involving visual content.

Read PDF

Similar papers

Open access Sep 2026

Evaluating the Accuracy of Large Language Models in Dentistry: A Multi-Model Study Using Clinical Questions From Turkey's Dental Specialty Exams.

AIM Artificial intelligence (AI) has become an integral part of dental education and clinical practice. While several studies have assessed the performance of ChatGPT in medical exams, comparative analyses of various large language models (LLMs) in dentistry remain scarce. This study aimed to evaluate and compare the p...

N. Acar, Fatih Sengul, Periş Çelikel et al. · 0 citations
Open access Aug 2026

Assessing the Educational Role of Large Language Models in Dental Training: A Decade-Long Analysis of Text-Based and Visual Questions in a National Examination.

LLMs demonstrate strong potential as supportive tools in dental education, particularly for text-based knowledge assessment, but their limited performance in visual question contexts highlights an important constraint for their integration into image-dependent domains such as dental diagnostics.

Sedef Ayşe Taşyapan, Didem Özer · 0 citations
Open access Aug 2026

Evaluating large language models in patient education: a comparative analysis addressing frequently asked questions in peri-acetabular osteotomy.

There may be significant differences in how effectively LLMs support patients with surgical queries, particularly in areas needing detailed explanation, and usually required minimal clarification in areas needing detailed explanation.

TP Davis, B. Guevel, K. Logishetty et al. · 0 citations
Open access Sep 2026

Assessment Design in the Era of Large Language Models: Evidence From Japanese Health Professions Licensing Examinations

Background The rapid development of large language models (LLMs) has raised questions about the continued educational value of conventional assessment formats in health professions education. This study evaluated how examination type, search access, image-based question characteristics, and item format influenced LLM a...

Toshitsugu Sakurai, Daichi Aizawa, Kazuyoshi Okawa et al. · 0 citations
Review Open access Jul 2026

Evaluating Large Language Models as Educational Tools: A Multi-Model Analysis of Performance on the Turkish Medical Specialty Examination (TUS) Ophthalmology Questions Özet:

This study aimed to comprehensively evaluate the performance of seven leading large language models (LLMs) from the 2024-2025 period on Turkish Medical Specialty Examination (TUS) ophthalmology questions, analyzing factors such as question type, chronology, and clinical area, while critically assessing the risk of data...

Burhan Başkan, Yusuf Evcimen, Özgür Bülent Timuçin et al. · 0 citations
Review Open access Sep 2026

Benchmarking Large Language Model Performance in Generating and Assessing Radiology Objective Structured Clinical Examination

High-quality radiology assessment questions are essential for education competency evaluation but labor-intensive to create. To compare four large language models (LLMs) in generating and evaluating radiology objective structured clinical examination (OSCE)–style questions and responses. Fifty Radio...

Ankush Ankush, Samriddhi Burman, Sydney Smith et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.