Cognitive complexity emerged as the strongest determinant of performance, with low-level questions significantly more likely to be answered correctly than high-level questions, and content domain associated with accuracy.
Abstract
Large language models (LLMs) are increasingly used to answer medical questions; however, their performance may vary depending on task characteristics. This study evaluated the performance of multiple versions of two widely used LLM families on oral and maxillofacial radiology (OMFR) questions from the Turkish Dental Specialty Examination (DUS) across three assessment phases and examined the influence of cognitive complexity, model family, evaluation phase, and content domain on response accuracy. A comparative repeated-evaluation design was used. A total of 123 text-based OMFR questions from DUS examinations (2012-2021) were submitted to two widely used LLM families (ChatGPT and DeepSeek) across three evaluation phases (May 2025, August 2025, and February 2026). Questions were categorized by content domain and Bloom cognitive level (low vs. high). Model responses were evaluated using official answer keys, and generalized estimating equations (GEE) were applied to account for repeated measurements. A total of 1230 model responses were analyzed, yielding an overall accuracy of 83.7%. Agreement between repeated runs was substantial to almost perfect (κ = 0.689-0.912). Cognitive complexity emerged as the strongest determinant of performance, with low-level questions significantly more likely to be answered correctly than high-level questions (OR = 6.15, p = 0.003). Content domain was also associated with accuracy (p = 0.028), whereas no statistically significant associations were observed for model family or evaluation phase. LLMs demonstrated high accuracy in answering OMFR examination questions; however, performance was more strongly associated with cognitive complexity than with model family or evaluation phase.
While LLMs show strong potential in supporting dental education through standardised exams, their performance varies by model and question type, and further improvements are needed to enhance reliability across different dental disciplines.
N. Acar, Fatih Sengul, Periş Çelikel et al.· European journal of dental e...· 0 citations
This study aimed to evaluate the performance of contemporary Large Language Models (LLMs) on the clinical sciences component of the Turkish Dental Specialization Examination (DUS) by comparing their accuracy across disciplines, examination years, and question types. The source dataset contained 1,040 scheduled clinical...
Fatih Karaaslan, Muhammed Halil Yilan, Merve Tirimoğulları· BMC Oral Health· 0 citations
Evaluating the reliability and readability of the responses generated by four mainstream LLMs to questions related to oral cancer found no model showed consistently high performance across all dimensions or met recommended readability standards.
Bo Zhang, Weidi Shi, Ying Zhang· Oral Health & Preventive Den...· 0 citations
LLMs demonstrate strong potential as supportive tools in dental education, particularly for text-based knowledge assessment, but their limited performance in visual question contexts highlights an important constraint for their integration into image-dependent domains such as dental diagnostics.
S. Taşyapan, Didem Özer· European journal of dental e...· 0 citations
All three large language models demonstrated high accuracy on operative dentistry MCQs, suggesting their potential utility as a supplementary educational tool, however, the residual error rates in the answers warrant the cautious integration of LLMs into dental education.
J. M. Antony, Nikhil Harikrishnan, N. Jayasheelan· Frontiers in Dental Medicine· 0 citations