Explainable Machine Learning for Diabetes Complication Prediction Using SHAP and Intervention Simulation
Abstract
The increasing availability of structured healthcare data has accelerated the use of Machine Learning (ML) to predict diabetes complications. However, limited interpretability remains a major barrier to clinical adoption. This study presents a unified framework that integrates predictive modeling, SHapley Additive exPlanations (SHAP), and constrained intervention simulation for interpretable multi-complication risk prediction in Type-2 Diabetes Mellitus (T2DM). A retrospective cohort study of 1011 individuals from a tertiary-care hospital in India was conducted to model the structure of clinical and lifestyle factors. Three algorithms, namely Random Forest (RF), XGBoost, and LightGBM, were developed to predict seven outcomes associated with diabetes complications. Model evaluation involved nested cross-validation, bootstrap confidence interval, precision and recall, calibration, F1-score, and Matthews Correlation Coefficient (MCC). The best predictive performance of Stroke was achieved with the RF algorithm, achieving an AUC-ROC of 0.957 (95% confidence interval: 0.901-0.993). Intervention simulation highlighted the predictive sensitivity of modifiable risk factors under clinically constrained feature modifications. The findings are predictive rather than causal and are limited by the single-center dataset without external validation.