Artificial Intelligence in Evaluating Undergraduate Dental Students: Benchmarking AI Platforms Against Student Performance.
Abstract
INTRODUCTION Artificial intelligence (AI), particularly large language models (LLMs), has demonstrated significant potential in educational assessment across medical disciplines. This study explores whether such models can match or surpass the performance of undergraduate dental students in a rigorous academic examination. This cross-sectional comparative study aims to compare the performance of three AI systems: ChatGPT, Gemini, and DeepSeek, with that of undergraduate dental students during a standardized final examination in prosthodontic technology.
Materials And Methods
A total of 49 second-year dental students completed a 45-item multiple-choice exam under standardized conditions. The same items were submitted to the AI models, and performance was analyzed using statistical methods including ANOVA, Cronbach Alpha and Tukey HSD to evaluate intergroup differences.
Discussion
The integration of generative AI into medical and dental education needs critical reevaluation of assessment integrity.
Results
DeepSeek achieved perfect scores on high-difficulty questions, while ChatGPT showed adaptive improvement across attempts. Both significantly outperformed the student cohort, whereas Gemini demonstrated lower and consistent accuracy; statistical analysis confirmed significant differences between models (P=0.036, η²=0.891).
Conclusions
AI models can equal or exceed student performance in domain-specific assessments, raising concerns about the integrity of unsupervised digital evaluations. Consequently, summative examinations should not be conducted online without strict proctoring or secure, verified environments to ensure academic validity and prevent unauthorized AI-assisted performance.