Skip to content
Open access

A Reliability-Aware and Interpretable Machine Learning Framework for Diabetes Prediction Using Structured Clinical Data

Aug 2026 · International Journal for Research in Applied Science and Engineering Technology · 0 citations

TL;DR

A reliability-aware and interpretable machine learning framework for diabetes prediction from structured clinical data is developed and a Feature Consistency Index (FCI) is formalised that quantifies the cross-model agreement of SHAP-derived feature importance and combines it with normalised importance into a single ranking score.

Abstract

Diabetes prediction plays an important role in re-ducing long-term health risks by enabling early medical interven-tion. Although machine learning models have been widely applied to this task, many existing studies emphasise predictive accuracy while giving comparatively little attention to the reliability, interpretability, and stability of the resulting decisions. This paper develops a reliability-aware and interpretable machine learning framework for diabetes prediction from structured clinical data. Three complementary models—Logistic Regression, Random Forest, and Extreme Gradient Boosting (XGBoost)—are trained on the Pima Indians Diabetes dataset so that both simple linear and complex non-linear relationships are captured. Beyond conventional discrimination metrics, the reliability of the predicted probabilities is quantified using the Brier score and reliability (calibration) diagrams. Interpretability is addressed with SHapley Additive exPlanations (SHAP) at both the global (cohort) and local (individual patient) levels. Because different models frequently emphasise different predictors, we formalise a Feature Consistency Index (FCI) that quantifies the cross-model agreement of SHAP-derived feature importance and combines it with normalised importance into a single ranking score. Finally, a perturbation-based robustness analysis measures the sensitivity of each model’s output to small changes in the input record. Experi-mentally, XGBoost achieves the highest discrimination (accuracy 0.7597, ROC-AUC 0.8374), whereas Random Forest attains the best-calibrated probabilities (Brier score 0.1646), demonstrating that discrimination and reliability are not interchangeable. The FCI identifies Glucose and BMI as simultaneously the most influential and the most consistently attributed predictors, while Blood Pressure and Skin Thickness are both weak and unstable. Under a 5% Gaussian perturbation of a representative patient record, the linear and bagged models shift by less than 0.01 in predicted probability, whereas the boosted model shifts by 0.0386, revealing an accuracy–stability trade-off that a purely accuracy-driven evaluation would not expose

Read PDF

Similar papers

Open access Jul 2026

An interpretable machine learning framework for early-stage diabetes mellitus prediction using comparative classification models and SHAP

A machine learning-based framework enhanced with explainability is introduced, built around a structured data preparation process that handles categorical encoding, numerical scaling, and minority class oversampling through the SMOTE technique, positioning it as a trustworthy tool for assisting medical professionals in...

N. J, Deekshitha U, K. V · 0 citations
Open access 2026

Diabetes onset prediction using random forest: A machine learning approach with the Pima Indians diabetes dataset

Diabetes, a chronic metabolic disorder, has affected millions of people worldwide, thus it is important to develop accurate predictive models for early intervention and improved patient prognosis. This paper aims to introduce a predictive model for diabetes onset using the Pima Indians Diabetes Dataset and the random f...

Jessie R. Paragas, Kent Claire Apple Joy M. Pallomina, Alexis Luke G. Barlomento · 0 citations
Open access Sep 2026

Explainable diabetes prediction using a stacked ensemble framework

Diabetes is a chronic disease that significantly increases the risk of serious complications such as cardiovascular disorders and kidney failure. Early detection through predictive modeling can lead to timely interventions and significantly improve patient health outcomes. Several machine learning approaches have been...

Saher Fatima Awan, Kainat Irfan, Umair Muneer Butt et al. · 0 citations
Conference Aug 2026

An Integrated Predictive, Explainable, and Causal Machine Learning Framework for Early Detection of Type 2 Diabetes with Uncertainty and Fairness Assessment

Early detection of Type 2 Diabetes Mellitus (T2DM remains a critical challenge in preventive healthcare due to the complex interplay of clinical, demographic, and lifestyle factors. This study proposes an Integrated Predictive, Explainable, and Causal Machine Learning Framework for early detection of Type 2 diabetes, i...

C. S. Reddy, Mohan Annamalai · 0 citations
Open access Aug 2026

Explainable Machine Learning for Diabetes Complication Prediction Using SHAP and Intervention Simulation

The increasing availability of structured healthcare data has accelerated the use of Machine Learning (ML) to predict diabetes complications. However, limited interpretability remains a major barrier to clinical adoption. This study presents a unified framework that integrates predictive modeling, SHapley Additive exPl...

R. U, M. P. Pushpalatha · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.