A comparative performance analysis of ensemble learning and regularized neural networks in cardiovascular risk prediction
Abstract
Introduction Heart disease classification using small structured clinical datasets requires evaluation procedures that account for sampling variability, model-selection bias, probability calibration, and clinically relevant operating characteristics. This study presents an internally validated comparison of classical classifiers, modern tree ensembles, and regularized artificial neural networks using a cleaned Cleveland heart disease dataset containing 303 observations, 13 predictors, and a binary outcome. Methods Continuous variables were standardized, nominal variables were one-hot encoded, and all preprocessing transformations were fitted exclusively within training folds. Model performance was evaluated using five-fold outer cross-validation repeated five times, with inner stratified cross-validation for hyperparameter optimisation. Twelve models were assessed using accuracy, sensitivity, specificity, positive and negative predictive values, F1-score, Matthew's correlation coefficient, ROC-AUC, PR-AUC, Brier score, and calibration measures. Results CatBoost achieved the highest mean ROC-AUC of 0.909 ± 0.041 and F1-score of 0.862 ± 0.041, while Logistic Regression achieved a comparable ROC-AUC of 0.906 ± 0.041. Random Forest obtained the highest mean PR-AUC of 0.916 ± 0.040. CatBoost did not significantly outperform Logistic Regression in ROC-AUC, PR-AUC, F1-score, or paired classification errors. Held-out permutation analysis identified the number of major vessels, thalassemia status, and chest-pain category as the most influential predictors. Discussion The findings indicate that no model family was universally superior and that a carefully regularized linear model remained competitive with substantially more complex approaches. These findings represent internal validation only and require confirmation using independent clinical cohorts before deployment.