Skip to content

KAISEN: Reproducible Subgroup Fairness Auditing for Clinical Risk Models

Jul 2026 · arXiv.org · Vol abs/2607.28608 · 0 citations · 49 references
Computer Science Biology

TL;DR

KAISEN, a five-phase audit pipeline covering subgroup stratification, disparity measurement, mechanism diagnostics, post-hoc mitigation, and drift monitoring, is evaluated to the point of failure on a synthetic benchmark of 16 disease tasks, 15 social-determinant axes from Healthy People 2030, and three prespecified intersections.

Abstract

Clinical risk models routinely achieve strong aggregate performance while producing materially different error rates across patient subgroups. Audit pipelines have been proposed to catch this, but their components are rarely stress-tested, so it is unclear which parts of an audit can be trusted and under what conditions. We present KAISEN, a five-phase audit pipeline covering subgroup stratification, disparity measurement, mechanism diagnostics, post-hoc mitigation, and drift monitoring, evaluated to the point of failure on a synthetic benchmark of 16 disease tasks, 15 social-determinant axes from Healthy People 2030, and three prespecified intersections. Four findings follow. (i) Significance tracks each axis's gap against its own minimum detectable effect: rank correlation between significance count and raw equalized-odds difference (EOD) across the 15 axes is rho = 0.56, rising to rho = 0.78 once EOD is standardized by that floor. (ii) Per-group threshold optimization reduces EOD in 48 of 48 held-out runs (paired delta = -0.285, 95% CI [-0.313, -0.252]), while group-wise Platt scaling -- the better calibrator -- behaves as a coin flip on EOD (19 of 48 runs improved, 95% CI [0.26, 0.55]) with mean effect near zero, so what an audit should report is the variance, not the average. (iii) The mechanism diagnostic classifies 144 of 144 controlled cases correctly but recovers none of 48 model-driven cases under proxy misspecification, with no signal that it failed. (iv) CUSUM failures and false alarms track cohort realization far more than disease: at the reference threshold, all 27 false alarms and 7 of 8 missed shifts come from different seeds (chi-squared p = 0.002), so a threshold tuned on one cohort fails to transfer. All results are synthetic with known ground truth and do not establish clinical validity. Code, artifacts, and scripts reproducing every number are released.

View source

Similar papers

Conference Aug 2026

Auditing Population-Level XAI Agreement with cABC: Evidence from Diabetes Risk Prediction

Auditing agreement between global explanation methods is underdeveloped in clinical XAI. Standard population-level agreement measures are poorly aligned with the practical question of interest: top-K overlap depends on an arbitrary cutoff, while rank-correlation metrics can overweight tail-order differences that are op...

V. Thieu, Hung-Nghiep Tran · 0 citations
#artificial intelligence Preprint Aug 2026

FairGlucose: A CGM Fairness Benchmark Reveals Subgroup Disparities Hidden in Population-Level Validation

It is established that population-level validation alone is insufficient for equity assessment of digital health AI, motivating subgroup-disaggregated reporting as a default standard, and subgroup-disaggregated reporting as a default standard for personalized configurations.

Junjie Luo, Xuzhe Zhi, Rui Han et al. · 0 citations
#machine learning Preprint Sep 2026

Counterfactual Fairness Audits of Multi-Step Clinical LLM Agents Require a Measured Per-Action Instability Floor

Counterfactual audits are the standard tool for checking whether a clinical agent treats demographically distinct but clinically identical patients differently. They report a flip rate: how often an action changes when only the patient descriptor changes. We show that this quantity is uninterpretable on its own. Re-run...

Rohith Reddy Bellibatlu, Manpreet Singh, Deepak Parashar et al. · 0 citations
Preprint Sep 2026

Differentially Private and Fairness-Audited Score Diffusion for Irregular Longitudinal Health Records

Sharing irregular longitudinal health records can accelerate model development, yet synthetic releases may leak participation, distort temporal dependence, suppress rare events, or reduce utility for underrepresented groups. We present TRUST LONGSYNTH, an auditable patient level private generator that combines bounded...

Taimoor Ahmad · 0 citations
Open access 2026

Severity-Weighted Calibration Error for Reliability Assessment Under Outcome-Severity Imbalance

Probability calibration is critical for reliable high-stakes predictive modelling, yet standard aggregate metrics such as Expected Calibration Error (ECE) can obscure important failures in severity-imbalanced subgroups. This work proposes and evaluates Severity-Weighted Calibration Error (SWCE), a formally defined post...

Aman Chandra H, K. Narendra · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.