Skip to content
Open access

COMPARATIVE ANALYSIS OF RANDOM FOREST AND SUPPORT VECTOR MACHINE ALGORITHMS FOR DIABETES MELLITUS PREDICTION

Aug 2026 · Jurnal Informatika dan Teknik Elektro Terapan · 0 citations

TL;DR

Comparisons of the performance of the Random Forest and Support Vector Machine algorithms in predicting diabetes and the effect of applying the Synthetic Minority Over-sampling Technique to imbalanced data show that Random Forest outperforms SVM.

Abstract

Abstract. Early prediction of type 2 diabetes mellitus is important to facilitate faster and more accurate diagnosis and clinical decision-making. This study aims to compare the performance of the Random Forest and Support Vector Machine (SVM) algorithms in predicting diabetes and to analyze the effect of applying the Synthetic Minority Over-sampling Technique (SMOTE) to imbalanced data. The study used the Bangladesh Diabetes 2025 dataset, following the stages of data selection, preprocessing, data transformation, modeling, and evaluation. The preprocessing stage included median imputation and the removal of duplicate data, while feature standardization was performed after data splitting to prevent data leakage. The dataset was split using an 80:20 ratio with a stratification technique. The study applied two experimental scenarios: one without SMOTE and one with SMOTE applied only to the training data. Model evaluation was conducted using the metrics accuracy, precision, recall, and F1-score based on the test data. The results show that Random Forest outperforms SVM. In the scenario without SMOTE, Random Forest achieved an accuracy of 94.8%, precision of 96.4%, recall of 97.0%, and an F1-score of 96.7%, while SVM achieved an accuracy of 91.5% and an F1-score of 94.6%. After applying SMOTE, Random Forest’s performance improved slightly to an accuracy of 95.3% and an F1-score of 97.0%, while SVM’s performance declined to an accuracy of 90.6% and an F1-score of 93.9%. The results of the study show that Random Forest is the best model for predicting type 2 diabetes mellitus in the dataset used. 

Read PDF

Similar papers

Open access Jul 2026

The Implementation of Support Vector Machine and Naïve Bayes Algorithm to Predict Diabetes

The experimental results show that for the GNB model, the best performance was achieved using the combination of StandardScaler, SMOTE, and SelectKBest (k=5), reaching an accuracy of 94.53%, precision 98.36%, recall 90.91%, and f1-score 94.49%.

Joshua Roy Danna Lacanlale, Vitri Tundjungsari · 0 citations
Open access Aug 2026

A Comparative Evaluation of Machine Learning Algorithms for Diabetes Risk Prediction

Evaluated machine learning algorithms for predicting diabetes risk from routinely available clinical and lifestyle variables confirm that ensemble tree-based methods, particularly Random Forest, provide a reliable, interpretable, and deployable basis for diabetes risk screening, especially in resource-constrained setti...

T. Olayinka · 0 citations
Open access Jul 2026

Predictive modeling of early diabetes diagnosis: An evaluation of XGBoost, support vector machine, and random forest classifiers

It is recommended that healthcare systems adopt XGBoost-based predictive models in clinical decision support tools for early screening, while future studies should validate these models using real-world clinical data to enhance reliability and generalizability.

Idehen Emmanuel Imafidon, Chikere Obinna Munachiso, Dominic Evans Onyebuchi et al. · 0 citations
Open access Aug 2026

Optimization of Diabetes Mellitus Classification Using the Random Forest and SMOTE-ENN Methods

Diabetes Mellitus is a non-communicable disease (NCD) that has turned into a worldwide health issue with a steadily rising prevalence. Timely identification is essential for minimizing the risk of complications and the financial strain on the healthcare system. This research focuses on creating a precise and dependable...

Imam Fadhur Rahman, Egia Rosi Subhiyakto, Cinantya Paramita · 0 citations
Open access 2026

Diabetes onset prediction using random forest: A machine learning approach with the Pima Indians diabetes dataset

Diabetes, a chronic metabolic disorder, has affected millions of people worldwide, thus it is important to develop accurate predictive models for early intervention and improved patient prognosis. This paper aims to introduce a predictive model for diabetes onset using the Pima Indians Diabetes Dataset and the random f...

Jessie R. Paragas, Kent Claire Apple Joy M. Pallomina, Alexis Luke G. Barlomento · 0 citations
Conference Aug 2026

Diabetes Disease Classification Using Tree-Based Machine Learning Algorithms: Random Forest, XGBoost, and LightGBM

The number of people affected by diabetes continues to grow worldwide, making it a major public health concern because delayed diagnosis can lead to severe health complications. This study applies and compares three treebased machine learning classifiers, namely Random Forest (RF), Extreme Gradient Boosting (XGBoost),...

Pragipta Septyaningrum Larasati, Hasih Pratiwi, Irwan Susanto et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.