Assessing the Educational Role of Large Language Models in Dental Training: A Decade-Long Analysis of Text-Based and Visual Questions in a National Examination.
Aug 2026· European journal of dental education· 0 citations· 11 references
Medicine
TL;DR
LLMs demonstrate strong potential as supportive tools in dental education, particularly for text-based knowledge assessment, but their limited performance in visual question contexts highlights an important constraint for their integration into image-dependent domains such as dental diagnostics.
Abstract
Background
Large language models (LLMs), including ChatGPT, Gemini and DeepSeek, are increasingly explored as supportive tools in health professions education. However, their educational utility across different knowledge domains and question formats, particularly those involving visual content, remains insufficiently understood.
Objective
This study aimed to evaluate the educational potential of three advanced LLMs in dental training by analysing their performance on a national specialty examination over a 10-year period, with particular emphasis on domain-specific accuracy and differences between text-based and visual questions.
Methods
A total of 1560 multiple-choice questions from 13 administrations of the Turkish Dental Specialty Examination (DUS) between 2012 and 2021 were included. All questions were translated into English and categorized into nine dental specialties. Each question was individually entered into ChatGPT-4.0, Gemini Advanced and DeepSeek in isolated sessions. Model responses were compared with official answer keys, and accuracy rates were analysed across years, specialties and question types. Statistical analyses included chi-square tests, one-way ANOVA and Kruskal-Wallis tests.
Results
ChatGPT-4.0 achieved the highest overall accuracy (86.93%), followed by Gemini (82.86%) and DeepSeek (82.34%). Performance varied across specialties, with higher accuracy observed in Basic Sciences and Oral Surgery, and lower performance in Orthodontics and Endodontics. A substantial decrease in accuracy was observed for visual questions (ChatGPT: 52.38%; Gemini: 45.24%; DeepSeek: 2.38%) compared to text-based items (all models > 83%). ChatGPT demonstrated more stable performance across years, whereas Gemini and DeepSeek showed greater variability.
Conclusions
LLMs demonstrate strong potential as supportive tools in dental education, particularly for text-based knowledge assessment. However, their limited performance in visual question contexts highlights an important constraint for their integration into image-dependent domains such as dental diagnostics. These findings underscore the need for cautious and context-aware implementation of AI tools in dental curricula and assessment practices.
AIM
Artificial intelligence (AI) has become an integral part of dental education and clinical practice. While several studies have assessed the performance of ChatGPT in medical exams, comparative analyses of various large language models (LLMs) in dentistry remain scarce. This study aimed to evaluate and compare the p...
N. Acar, Fatih Sengul, Periş Çelikel et al.· European journal of dental e...· 0 citations
AI models demonstrate generally high accuracy in medical examinations; however, they struggle with interpreting clinical context, recognising atypical medical scenarios and answering questions involving visual content.
Bünyamin Arı· Health Informatics Journal· 0 citations
This study aimed to comprehensively evaluate the performance of seven leading large language models (LLMs) from the 2024-2025 period on Turkish Medical Specialty Examination (TUS) ophthalmology questions, analyzing factors such as question type, chronology, and clinical area, while critically assessing the risk of data...
Burhan Başkan, Yusuf Evcimen, Özgür Bülent Timuçin et al.· Muş Alparslan Üniversitesi S...· 0 citations
Background The rapid development of large language models (LLMs) has raised questions about the continued educational value of conventional assessment formats in health professions education. This study evaluated how examination type, search access, image-based question characteristics, and item format influenced LLM a...
Toshitsugu Sakurai, Daichi Aizawa, Kazuyoshi Okawa et al.· Journal of Medical Education...· 0 citations
This study aimed to evaluate the performance of contemporary Large Language Models (LLMs) on the clinical sciences component of the Turkish Dental Specialization Examination (DUS) by comparing their accuracy across disciplines, examination years, and question types.
The source dataset contained 1,040 sched...
Fatih Karaaslan, Muhammed Halil Yılan, Merve Tirimoğulları· BMC Oral Health· 0 citations
Background: Large language models (LLMs) have demonstrated strong performance on standardized medical examinations, with recent studies reporting performance approaching or exceeding that of senior medical residents. However, examination accuracy alone does not establish how models arrive at their answers or the relati...
F. Gafoor, M. Syed, M. Halai et al.· medRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.