Assessing the Diagnostic Performance of ChatGPT-5.0 versus Machine Learning in Orthodontics: A Comparative Analysis for Extraction Treatment Planning.
Abstract
Objective To make accurate orthodontic extraction decisions, various clinical and cephalometric variables must be evaluated. This study aims to evaluate ChatGPT-5.0's performance in distinguishing orthodontic extraction decisions and to compare it with five supervised machine learning (ML) algorithms. Methods Of 550 retrospectively evaluated orthodontic records, 30 were reserved for calibration, leaving 520 for the main analysis. The reference standard was the consensus treatment decision of three expert orthodontists with more than 5 years of clinical experience. Overall, 23 variables were analyzed, including 13 clinical parameters, 7 cephalometric measurements, and photographs. ChatGPT-5.0's performance was evaluated using a 5-fold cross-validation design. It was compared with XGBoost, random forest, support vector machine (SVM), logistic regression, and multi-layer perceptron (MLP). Performance metrics included accuracy, sensitivity, specificity, precision, F1-score, and balanced accuracy, with 95% confidence intervals calculated. Statistical analyses utilized Cochran's Q test and the McNemar test with Holm-Bonferroni correction. Results Of the 520 main cases, 223 (42.88%) were extraction treatments and 297 (57.12%) were non-extraction treatments. XGBoost achieved the highest accuracy (78.08%), followed closely by ChatGPT-5.0 (75.77%). The overall performance difference among models was significant (p≤0.001). In pairwise comparisons, ChatGPT's accuracy was significantly higher than those of random forest, SVM, logistic regression, and MLP, but was found to be similar to XGBoost. ChatGPT-5.0 showed the highest sensitivity (76.68%), whereas XGBoost showed the highest specificity (82.15%). Conclusion ChatGPT-5.0 demonstrated performance comparable to, and in some cases superior to, traditional ML models for orthodontic extraction decisions. While XGBoost yielded the highest overall classification accuracy, ChatGPT-5.0's high sensitivity in detecting extraction cases was noteworthy.