KAISEN, a five-phase audit pipeline covering subgroup stratification, disparity measurement, mechanism diagnostics, post-hoc mitigation, and drift monitoring, is evaluated to the point of failure on a synthetic benchmark of 16 disease tasks, 15 social-determinant axes from Healthy People 2030, and three prespecified intersections.
Abstract
Clinical risk models routinely achieve strong aggregate performance while producing materially different error rates across patient subgroups. Audit pipelines have been proposed to catch this, but their components are rarely stress-tested, so it is unclear which parts of an audit can be trusted and under what conditions. We present KAISEN, a five-phase audit pipeline covering subgroup stratification, disparity measurement, mechanism diagnostics, post-hoc mitigation, and drift monitoring, evaluated to the point of failure on a synthetic benchmark of 16 disease tasks, 15 social-determinant axes from Healthy People 2030, and three prespecified intersections. Four findings follow. (i) Significance tracks each axis's gap against its own minimum detectable effect: rank correlation between significance count and raw equalized-odds difference (EOD) across the 15 axes is rho = 0.56, rising to rho = 0.78 once EOD is standardized by that floor. (ii) Per-group threshold optimization reduces EOD in 48 of 48 held-out runs (paired delta = -0.285, 95% CI [-0.313, -0.252]), while group-wise Platt scaling -- the better calibrator -- behaves as a coin flip on EOD (19 of 48 runs improved, 95% CI [0.26, 0.55]) with mean effect near zero, so what an audit should report is the variance, not the average. (iii) The mechanism diagnostic classifies 144 of 144 controlled cases correctly but recovers none of 48 model-driven cases under proxy misspecification, with no signal that it failed. (iv) CUSUM failures and false alarms track cohort realization far more than disease: at the reference threshold, all 27 false alarms and 7 of 8 missed shifts come from different seeds (chi-squared p = 0.002), so a threshold tuned on one cohort fails to transfer. All results are synthetic with known ground truth and do not establish clinical validity. Code, artifacts, and scripts reproducing every number are released.
Auditing agreement between global explanation methods is underdeveloped in clinical XAI. Standard population-level agreement measures are poorly aligned with the practical question of interest: top-K overlap depends on an arbitrary cutoff, while rank-correlation metrics can overweight tail-order differences that are op...
V. Thieu, Hung-Nghiep Tran· International Conference on...· 0 citations
It is established that population-level validation alone is insufficient for equity assessment of digital health AI, motivating subgroup-disaggregated reporting as a default standard, and subgroup-disaggregated reporting as a default standard for personalized configurations.
Junjie Luo, Xuzhe Zhi, Rui Han et al.· 0 citations
Counterfactual audits are the standard tool for checking whether a clinical agent treats demographically distinct but clinically identical patients differently. They report a flip rate: how often an action changes when only the patient descriptor changes. We show that this quantity is uninterpretable on its own. Re-run...
Sharing irregular longitudinal health records can accelerate model development, yet synthetic releases may leak participation, distort temporal dependence, suppress rare events, or reduce utility for underrepresented groups. We present TRUST LONGSYNTH, an auditable patient level private generator that combines bounded...
Probability calibration is critical for reliable high-stakes predictive modelling, yet standard aggregate metrics such as Expected Calibration Error (ECE) can obscure important failures in severity-imbalanced subgroups. This work proposes and evaluates Severity-Weighted Calibration Error (SWCE), a formally defined post...
Aman Chandra H, K. Narendra· IEEE Access· 0 citations
VFR-Audit is proposed, a framework built around the Verdict Flip Rate (VFR), a scalar bounded between 0 and 0.5 that measures the probability of verdict reversal under stratified bootstrap resampling.
Merin Joy, V. Võ, Caslon Chua· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.