Enhancing Multiclass Diabetes Prediction with SMOTE and Machine Learning Classifiers
Abstract
Proper identification of the stages of diabetes (non-diabetic, prediabetic, diabetic) is pivotal in the early intervention and risk stratification. Nevertheless, class imbalance in clinical datasets often imbalances machine learning models in the context of majority classes. In this paper, we have assessed four classifiers namely, the Support Vector Machine (SVM), Artificial Neural Network (ANN) and the Logistic Regression (LR) and on a multiclass dataset on diabetes. Once the duplicate cases of patients were eliminated 264 distinct cases were retained. SMOTE was only used on the training data to overcome the problem of class imbalance. Findings indicate that Random Forest performed the highest with an accuracy of 98.1, F1-score of 98.5, and AUC of 99.9 and stability before and after balancing. Logistic Regression and ANN demonstrated some significant improvements in recall following SMOTE, which illustrates the costs and benefits of operating a global model versus a minority-class sensitive one. SHAP analysis found that HbA1c, age, LDL, and BMI are the most influential predictors that improve interpretability and clinical relevance. These results prove that ensemble models applied to class balancing and explainability methods can be used to rely on and provide clear tools regarding multiclass diabetes classification.