Comparative performance and temporal variability of large language models on orthodontic questions from a national dental specialty examination
This study compared the performance of three artificial intelligence–based chatbots (ChatGPT-4o, ChatGPT-4.5, and Gemini 2.5 Pro) on orthodontic questions from the Turkish Dental Specialty Examination at two testing time points. A total of 179 orthodontic multiple-choice questions from 18 examinations conducted between...