Accuracy of large language models in the Turkish dental specialization examination (DUS): a multidimensional evaluation across disciplines and question formats
This study aimed to evaluate the performance of contemporary Large Language Models (LLMs) on the clinical sciences component of the Turkish Dental Specialization Examination (DUS) by comparing their accuracy across disciplines, examination years, and question types. The source dataset contained 1,040 sched...