Aug 2026· Journal of Intelligent Decision Making and Information Science· Vol 3, pp. 273-294· 0 citations· 29 references
TL;DR
The pipeline approach ensured no data leakage in cross-validation, and the findings support ensemble ML models with SMOTE as a preprocessing step for imbalanced CVD datasets.
Abstract
Cardiovascular disease (CVD) is the major cause of global mortality, accounting for approximately 17.9 million annual deaths and 32% of all global fatalities. The dataset revealed a real class imbalance of 58% CVD-positive cases (580) and 42% CVD-negative cases (420), with an imbalance ratio of 1.381. Without correction, classifiers trained on this data are biased toward the majority class, inflating accuracy while suppressing sensitivity on the minority class. To correct class imbalance using the Synthetic Minority Over-sampling Technique (SMOTE) on the training set, compare six supervised machine learning classifiers for binary CVD detection, and evaluate their performance on the original test set. Six classifiers Decision Tree, KNN, SVM, Gradient Boosting, Random Forest, and Logistic Regression were evaluated. A stratified 80:20 train-test split was applied, followed by SMOTE only on the training set (336 → 464 samples per class). To prevent data leakage, SMOTE was embedded within an imbalanced-learn Pipeline. Performance on the original test distribution was evaluated using Accuracy, Precision, Recall, F1 Score, MCC, and AUC-ROC. After SMOTE balancing, Logistic Regression and Random Forest achieved the highest F1 score (0.9871), with AUC-ROC values of 0.9982 and 0.9990, respectively. Gradient Boosting achieved the highest AUC-ROC (0.9993) and was the most stable model in cross-validation (99.20% ± 0.98%). Feature importance analysis identified maxheartrate, oldpeak, and noofmajorvessels as the top three discriminative predictors. The pipeline approach ensured no data leakage in cross-validation, and the findings support ensemble ML models with SMOTE as a preprocessing step for imbalanced CVD datasets.
Proper identification of the stages of diabetes (non-diabetic, prediabetic, diabetic) is pivotal in the early intervention and risk stratification. Nevertheless, class imbalance in clinical datasets often imbalances machine learning models in the context of majority classes. In this paper, we have assessed four classif...
M. H. Ahmed, J. Qadir· Academic Journal of Internat...· 0 citations
Cardiovascular disease prediction using structured clinical data is commonly formulated as a binary classification problem, while multi-class prediction is more challenging because of overlapping clinical characteristics and class imbalance. This study presents a hybrid ensemble and imbalance-aware machine-learning fra...
The study comes to the conclusion that headline accuracy is an unreliable guide in imbalanced medical prediction, that imbalance handling can change a model's practical usefulness, and that this benefit is strongly algorithm-dependent, meaning that the decision to resample should be based on the algorithm and the scree...
A. Oduroye, Temilade Opanuga, Esther Tosin Akanbi et al.· International journal of re...· 0 citations
Cardiovascular diseases remain a leading cause of mortality worldwide. Identifying underlying clinical phenotypes early, such as distinct categories of chest pain, is vital for diagnostic triage and downstream medical decision-making. This study evaluates the performance of four prominent machine learning algorithms –...
Twana Abdulqader Mohammed, S. Salh· UHD Journal of Science and T...· 0 citations
Heart disease is a leading cause of mortality worldwide, with early detection playing a critical role inreducing death rates. Accurate prediction of heart disease remains challenging due to complex medical data andthe inability to provide continuous monitoring. Utilizing the Heart Disease dataset, various feature selec...
Manoj Kumar Konudula, S. K, R. M· Advanced International Journ...· 0 citations
Experimental results demonstrate that the optimized XGBoost-SMOTE model significantly outperforms traditional machine learning algorithms including Logistic Regression, Decision Tree, Random Forest, Support Vector Machine, KNearest Neighbors, AdaBoost, and baseline XGBoost.
B. Naveen, N. Rao· International Journal of Eng...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.