A Dual Explainability Framework for Cardiovascular Disease Prediction Using SHAP and LIME with Tree-Based Machine Learning Models
Abstract
Cardiovascular disease (CVD) remains one of the main cause of mortality worldwide and although ensemble learning models such as Random Forest, XGBoost, and LightGBM offers strong predictive performance, their black box nature limits adoption in clinical settings. This study proposes a dual explainability framework that integrate SHAP and LIME with tree based ensemble models for cardiovascular disease prediction using the Heart Failure Prediction Dataset, which consists of 918 patient records with eleven clinical features. Random Forest, XGBoost, and LightGBM were trained and optimized using GridSearchCV with five fold cross validation, with recall prioritized as the primary metric to minimize false negatives in cardiovascular risk screening. Random Forest achieved the highest cross validation recall of 0.9112 and selected as the best performing model, reaching a test accuracy of 0.85 and recall of 0.91 on the Heart Disease class. SHAP TreeExplainer then apply to providing global feature importances, identifying ST Slope, Chest Pain Type, Cholesterol, and Oldpeak as the most influential predictor, while LIME Tabular Explainer provides patient level explanation for individual prediction outcome. Result show that SHAP and LIME produced consistent explanation for correctly classified cases, aligning with established cardiovascular risk factor, demonstrating that high predictive performances and clinical interpretability can be achieved together.