Skip to content
Open access

A Comparative Evaluation of Machine Learning Algorithms for Diabetes Risk Prediction

Aug 2026 · FUDMA Journal of Sciences · 0 citations · 5 references

TL;DR

Evaluated machine learning algorithms for predicting diabetes risk from routinely available clinical and lifestyle variables confirm that ensemble tree-based methods, particularly Random Forest, provide a reliable, interpretable, and deployable basis for diabetes risk screening, especially in resource-constrained settings.

Abstract

Diabetes mellitus is a chronic metabolic disorder whose global prevalence continues to rise, creating an urgent need for scalable, low-cost tools for early risk identification. This study evaluates the effectiveness of five machine learning algorithms (Random Forest, XGBoost, Support Vector Machine [SVM], CatBoost, and TabNet) for predicting diabetes risk from routinely available clinical and lifestyle variables. Using the Pima Indians Diabetes dataset, a preprocessing pipeline was applied that included median imputation of physiologically implausible zero values, standardized (Z-score) feature scaling, Boruta-based feature selection, and class-imbalance handling through class-weight adjustment and the Synthetic Minority Oversampling Technique (SMOTE). Models were trained on an 80/20 train-test split and assessed using accuracy, precision, recall, F1-score, and the area under the receiver operating characteristic curve (ROC-AUC). Random Forest achieved the strongest overall performance (accuracy 0.753; F1-score 0.689; ROC-AUC 0.810), followed by CatBoost, SVM, and XGBoost, whereas TabNet performed worst with very low recall for the diabetic class. The best-performing model (Random Forest) was deployed in a lightweight Flask web application that returns a probability-based diabetes risk assessment, categorising each prediction as low, moderate, or high risk together with a tailored recommendation. The findings confirm that ensemble tree-based methods, particularly Random Forest, provide a reliable, interpretable, and deployable basis for diabetes risk screening, especially in resource-constrained settings. Key limitations include dataset homogeneity, residual class imbalance, and limited feature coverage.

Read PDF

Similar papers

Review Open access Jul 2026

A Comparative Evaluation of Three-Class Diabetes Classification Using Machine Learning Algorithms

Random Forest provided the best overall performance for three-class diabetes classification in this analytical sample, however, modest agreement, low prediabetes sensitivity, potential label leakage from fasting glucose, and the absence of external validation indicate that further evaluation is required before clinical...

Ayeni Taiwo Michael, Odukoya Ayooluwa, Ilesanmi Opeyemi · 0 citations
Open access Jul 2026

Predictive modeling of early diabetes diagnosis: An evaluation of XGBoost, support vector machine, and random forest classifiers

It is recommended that healthcare systems adopt XGBoost-based predictive models in clinical decision support tools for early screening, while future studies should validate these models using real-world clinical data to enhance reliability and generalizability.

Idehen Emmanuel Imafidon, Chikere Obinna Munachiso, Dominic Evans Onyebuchi et al. · 0 citations
Open access 2026

A Comparative Evaluation of Various Machine Learning Techniques for Prediction of Type 2 Diabetes Mellitus

This research uses Machine Learning algorithms to evaluate the likelihood of Type 2 Diabetes Mellitus (T2DM) using lifestyle and family history data and presents a performance assessment of seven ML classifiers, showing that SVM achieved the highest accuracy and AUC, while LR, RF, and XGBoost also performed competitive...

Rizwan Akhtar, Muhammad Kalamuddin Ahamad · 0 citations
Open access 2026

Machine learning approaches for diabetes prediction: A comparative analysis of classification models

Early and accurate identification of individuals at risk for type 2 diabetes mellitus (T2DM) is a clinical priority given the global scale of the epidemic. Machine learning (ML) methods offer a data-driven complement to conventional screening approaches; however, systematic comparisons of multiple classifiers under ide...

F. Yagin · 0 citations
Open access 2026

A Comparative Analysis of Machine Learning Algorithms for the Early Prediction of Diabetes with an Evaluation of Class-Imbalance Handling

The study comes to the conclusion that headline accuracy is an unreliable guide in imbalanced medical prediction, that imbalance handling can change a model's practical usefulness, and that this benefit is strongly algorithm-dependent, meaning that the decision to resample should be based on the algorithm and the scree...

A. Oduroye, Temilade Opanuga, Esther Tosin Akanbi et al. · 0 citations

Diabetes Prediction Using Machine Learning Model: A comparative Approach

Six supervised learning models were developed and compared for diabetes prediction using a dataset and compared for diabetes prediction using a 100k patients records with eight clinical features including gender, age, hypertension, smoking history, heart disease, BMI, HbA1c level, and blood glucose level.

Akshay Bhardwaj, Rajesh Chauhan, Devansh Khajuria · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.