Skip to content
Open access

Accuracy of large language models in the Turkish dental specialization examination (DUS): a multidimensional evaluation across disciplines and question formats

Sep 2026 · BMC Oral Health · 0 citations

Abstract

This study aimed to evaluate the performance of contemporary Large Language Models (LLMs) on the clinical sciences component of the Turkish Dental Specialization Examination (DUS) by comparing their accuracy across disciplines, examination years, and question types. The source dataset contained 1,040 scheduled clinical-science questions from the official DUS examinations (2012/1–2021). Thirteen officially cancelled items were excluded, leaving 1,027 evaluable questions across eight disciplines (871 MCQs, 119 CMCQs, and 37 IBMCQs). The same questions were submitted once to seven consumer AI interfaces. Overall and exploratory subgroup comparisons were performed using Cochran’s Q test. Only the overall comparison was followed by 21 exact pairwise McNemar tests with Holm adjustment. Wilson 95% confidence intervals were calculated for overall accuracy. Overall accuracy differed among the seven interfaces (Cochran’s Q = 669.10; p  < 0.001). Gemini 2.5 Pro achieved the highest accuracy (92.9%; 954/1,027), whereas Qwen3-Max achieved the lowest (59.5%; 611/1,027). Exploratory Cochran’s Q tests indicated omnibus differences among interfaces within every dental discipline (all p  < 0.001) and within MCQ (Q = 669.91; p  < 0.001), CMCQ (Q = 59.29; p  < 0.001), and IBMCQ groups (Q = 14.68; p  = 0.023). The IBMCQ result should be interpreted cautiously because only 37 items were available. Under standardized, single-attempt Turkish DUS testing conditions, the evaluated consumer AI interfaces showed materially different accuracy across the full dataset and exploratory discipline- and format-based subgroups. These findings characterize examination performance only and require confirmation with novel, unpublished, clinically contextualized, and independently assessed questions before broader educational or clinical utility can be inferred.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.