Imbalance-Aware Evaluation of Stroke Risk Prediction: A Methodological Benchmark Using Calibration and Decision Curve
Abstract
Stroke prediction models are often evaluated using accuracy and ROC-AUC, although these metrics can be misleading when the outcome is rare. This study presents a methodological benchmark—not a clinical validation study—for imbalance-aware evaluation of stroke risk prediction, emphasizing probability reliability, threshold performance, and decision-analytic utility. Logistic Regression (LR), Random Forest (RF), and Gradient Boosting (GB) were evaluated using the Kaggle Stroke Prediction Dataset (n = 5,110; 249 stroke cases; prevalence = 4.87%). Training folds were balanced by random oversampling, and an analytical prior correction was applied to adjust the class-prior shift introduced by oversampling. Evaluation included ROC-AUC, precision-recall AUC, Brier score, expected calibration error, calibration slope and intercept, exploratory F2-optimal thresholds, and decision curve analysis. GB achieved the highest discrimination (AUC = 0.8415; PR-AUC = 0.2126), whereas LR showed the best calibration (ECE = 0.0048; slope = 0.967). RF produced high accuracy but very low sensitivity and severe miscalibration (slope = 0.556), making its probabilities unsuitable for threshold-based interpretation without further recalibration. LR and GB generated comparable net benefit within a prespecified exploratory threshold range of 1–10%. These findings show that calibration-aware and decision-analytic evaluation is necessary for imbalanced prediction tasks. Because the dataset has uncertain clinical provenance and no external validation was performed, the results are not intended for direct clinical implementation.