Skip to content
Open access

Severity-Weighted Calibration Error for Reliability Assessment Under Outcome-Severity Imbalance

2026 · IEEE Access · Vol 14, pp. 116729-116743 · 0 citations · 44 references
Computer Science

Abstract

Probability calibration is critical for reliable high-stakes predictive modelling, yet standard aggregate metrics such as Expected Calibration Error (ECE) can obscure important failures in severity-imbalanced subgroups. This work proposes and evaluates Severity-Weighted Calibration Error (SWCE), a formally defined post-hoc calibration framework incorporating outcome-severity weighting, and demonstrates that aggregate evaluation — regardless of weighting scheme—cannot reliably reveal subgroup-specific miscalibration. This result motivates severity-band-stratified calibration gap analysis as the primary diagnostic framework for assessing reliability under outcome-severity imbalance. Five representative models are evaluated on MIMIC-IV (Medical Information Mart for Intensive Care, version 4; 72,001 patients) with temporal replication on MIMIC-III (45,278 patients) using a strict 70/10/20 stratified split and Sequential Organ Failure Assessment (SOFA)-defined severity bands, with no test-set leakage during threshold selection or recalibration. A calibration paradox is identified: the model with the best aggregate ECE simultaneously produces the most dangerous severe-band calibration gap, systematically underestimating mortality in the highest-acuity subgroup, while models with comparably or substantially worse aggregate calibration produce negligible undertriage of severe patients. This divergence between aggregate rank and subgroup safety is consistent across both datasets and across all tested classification thresholds. SWCE produces aggregate rankings identical to standard ECE, formally establishing that aggregate severity weighting cannot resolve the subgroup masking problem. Standard global recalibration is often ineffective, whereas validation-based isotonic recalibration within severity bands substantially reduces subgroup miscalibration on both datasets. These findings establish severity-stratified calibration evaluation as a practical requirement for reliable deployment of high-stakes predictive models.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.