Skip to content
Conference

Auditing Population-Level XAI Agreement with cABC: Evidence from Diabetes Risk Prediction

Aug 2026 · International Conference on Multimedia Analysis and Pattern Recognition · pp. 442-447 · 0 citations · 16 references

Abstract

Auditing agreement between global explanation methods is underdeveloped in clinical XAI. Standard population-level agreement measures are poorly aligned with the practical question of interest: top-K overlap depends on an arbitrary cutoff, while rank-correlation metrics can overweight tail-order differences that are operationally negligible. We address this problem by combining computed ABC (cABC) analysis with per-group Jaccard indices to compare feature-importance rankings at the level of data-driven importance groups rather than raw rank positions. We apply this framework to SHAP and Permutation Importance (PI) in diabetes risk prediction using three Behavioral Risk Factor Surveillance System (BRFSS) cohorts (2015: n = 253,680; 2021: n = 236,378; 2023: n = 272,769) and three classifiers: XGBoost, Random Forest, and Logistic Regression. Across nine model–cohort settings, cross-method agreement between SHAP and PI is high, with mean Group-A Jaccard JA =0.9127, whereas cross-model agreement is lower, with mean JA =0.8307. This suggests that, in our experimental setting, the choice of model pipeline may perturb global feature-importance structure more than attribution-method choice. Five features (Age, BMI, GenHlth, HighBP, and HighChol) remain essential across all 18 attribution rankings. These results position cABC-based group-overlap analysis as a data-driven framework for population-level XAI auditing. In this setting, the framework reveals a practically relevant observation for diabetes risk prediction: stable identification of essential predictors is more sensitive to model-pipeline choice than to whether explanations are generated with SHAP or PI. The source code and dataset are available at https://github.com/thieuanhvan/diabetes-xai-agreement.

View source

Similar papers

Preprint Sep 2026

A statistical framework for identifying subgroup vulnerability to predictive multiplicity in clinical AI

AI models trained on the same data can disagree about patient risk, with disagreement potentially concentrated in clinically important subgroups. We propose V(S), a statistically grounded vulnerability index combining an observable lower-bound witness of model disagreement with clinical severity, and develop inference...

Enock Adu Bonsu · 0 citations
#artificial intelligence Preprint Aug 2026

FairGlucose: A CGM Fairness Benchmark Reveals Subgroup Disparities Hidden in Population-Level Validation

It is established that population-level validation alone is insufficient for equity assessment of digital health AI, motivating subgroup-disaggregated reporting as a default standard, and subgroup-disaggregated reporting as a default standard for personalized configurations.

Junjie Luo, Xuzhe Zhi, Rui Han et al. · 0 citations
Open access Sep 2026

Towards Interpretable Risk: Multidimensional Context for ICU Mortality Predictions

ICU mortality models can achieve strong discrimination, yet a risk score alone provides limited context for patient-level interpretation. We developed a multidimensional prediction-context framework that complements a calibrated mortality estimate with model behavior, data availability, recent physiology, and model att...

S. Gupta, A. Das, M. S. Anto et al. · 0 citations
Review Open access Sep 2026

Development and validation of a prior-to-admission medication list risk scoring tool.

PURPOSE To develop, validate, and implement an admission predictive model that estimates the likelihood and expected number of changes to the prior-to-admission (PTA) medication list, to help triage pharmacy-led medication histories. METHODS We performed a retrospective study of adult admissions at a large academic m...

S. Nelson, M. Hobensack, L. Fleenor et al. · 0 citations
Review Open access Aug 2026

Counterfactual Analysis of Executable Clinical Decision Logic

Clinical recommendations are often expressed in narrative form, which limits their direct execution, auditability, and patient-specific interpretation. This paper presents a hybrid decision-support framework that combines Decision Model and Notation (DMN), survey-weighted rule-ensemble learning, and counterfactual sens...

C. Maleki, Y. Bertrand, F. Gailly · 0 citations
Review Open access Sep 2026

Multidimensional Validation of Clinical Evidence: A Four-Domain Audit Framework

INTRODUCTIONThe evaluation of clinical evidence continues to rely predominantly on measures of statistical significance and relative effect size. Although indispensable, these metrics alone may not adequately characterize evidence quality, particularly with respect to follow-up integrity, clinical relevance, predictive...

F. Masedu, Monica Mazza, M. Attanasio et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.