Skip to content
Preprint

A Statistical Framework for Data-Driven Discovery of Differential Performance in Clinical Risk Prediction Models

Aug 2026 · 0 citations · 16 references
Mathematics

TL;DR

The unfairness tree (utree) is proposed, a data-driven recursive partitioning framework for identifying subgroups with differential model performance that exhibits nominal empirical type I error rates and good ability to detect, quantify, and characterize performance discrepancies defined by higher-order variable interactions.

Abstract

Predictive models employing artificial intelligence (AI) and machine learning (ML) are increasingly being used for decision support in healthcare settings. These models may exhibit differential performance across population subgroups defined by race, age, sex, and other factors and cause disparate clinical impacts, leading to intensive recent study of what has been termed"model fairness". While many methods have been proposed to assess risk prediction model fairness, these techniques generally require that the end user pre-specify the groups across which fairness is to be evaluated. In real-world settings, however, important model performance disparities may arise in unknown subgroups defined by multiple intersecting characteristics. To address this problem, we propose the unfairness tree (utree), a data-driven recursive partitioning framework for identifying subgroups with differential model performance. In simulations, the utree exhibits nominal empirical type I error rates and good ability to detect, quantify, and characterize performance discrepancies defined by higher-order variable interactions. In six mortality risk models fit to the GUSTO-I acute myocardial infarction trial dataset, utrees identified subgroup-specific performance patterns, with age, sex, blood pressure, and Killip class consistently associated with differential model performance.

View source

Similar papers

Open access

A framework for fair and robust clinical risk prediction through collaborative learning and localized uncertainty quantification

Clinical prediction models trained on heterogeneous populations may exhibit disparities in predictive performance across patient subgroups, potentially reflecting differences in data availability, feature completeness, and representation in training data. Existing fairness-aware approaches can introduce performance-fairness trade-offs and commonly optimize group-level disparities without explicitly accounting for their subgroup-specific sources, limiting both their effectiveness and their clinical interpretability. This dissertation makes two methodological contributions. The first is a collaborative learning framework that treats demographic subgroups as clients and aggregates their model parameters using strategies that encode different assumptions about group contribution. Rather than optimizing an explicit group-fairness constraint, the framework preserves and integrates subgroup-specific information to improve equity while largely maintaining subgroup predictive performance. The second contribution is the localized conformal classifier, a neighborhood-adaptive conformal prediction framework for instance-level uncertainty quantification. By calibrating prediction sets according to the density and composition of a patient's local neighborhood in the feature space, the method communicates both the model's prediction and the uncertainty associated with that individual prediction. This information is particularly consequential when predictions inform clinical decision-making. Both contributions are evaluated across three clinical prediction tasks: 30-day mortality in ICU patients with sepsis, treatment non-completion in patients with substance use disorder, and hospital readmission after surgical resection in patients with colorectal cancer. Compared with reweighting and constrained optimization, collaborative learning achieves more favorable fairness-performance trade-offs with less degradation in predictive performance, and its variants consistently appear on the Pareto frontier across all three prediction tasks. A novel difficulty decomposition framework, derived from the localized conformal classifier's calibration scores, distinguishes sources of uncertainty arising from the local neighborhood from those specific to an individual prediction. SHAP attribution reveals how collaborative learning and fairness interventions alter the feature-attribution patterns in the underlying prediction models. Surrogate regression trees trained on the localized conformal classifier's difficulty scores relate neighborhood and instance difficulty to clinical features observable at the point of care. Clinical features remain strongly associated with instance-level difficulty under collaborative learning, whereas these associations weaken substantially after reweighting and are absent in some analyses. Across all three clinical contexts, local outcome heterogeneity is the most consistent driver of neighborhood-level predictive difficulty, while proximity to the decision boundary is the principal driver of instance-level difficulty. Together, these contributions provide a framework for developing and evaluating clinical prediction models that are not only accurate, but also more equitable across patient groups and more informative about the reliability of individual predictions.

Mary M. Lucas, Christopher C. Yang · 0 citations
Open access Aug 2026

Comparison of Methods for Incorporating Related Data When Developing Clinical Prediction Models: A Simulation Study

ABSTRACT Clinical Prediction Models (CPMs) compute an individual's risk of an outcome, given a set of predictors. Guidance states CPMs should be constructed using data sampled from the target population. However, researchers might have access to ancillary data sets from different time points, countries, or healthcare settings, which could support model development. This study explores in which situations the ancillary data affect CPM performance, given potential heterogeneity. We conducted a simulation study to assess the impact of heterogeneity between the target and the ancillary data sets, and their relative sample size, on CPM performance. Target and ancillary populations were generated with varying degrees of heterogeneity. CPMs were developed using target‐only logistic regression, logistic and intercept updating, and importance weighting using propensity scores. These models were evaluated on independent data using calibration, discrimination, and prediction stability. Also, a real‐world case study was used as an illustrative example of application using the SWEDEHEART registry. Incorporating ancillary data generally improve CPM performance. Logistic and Intercept Recalibration often outperformed the target‐only regression approach. However, Logistic Recalibration showed greater variability and instability in calibration, while Intercept Recalibration performed poorly under predictor–outcome association shift. The importance weighting method demonstrated consistent performance across a wide range of scenarios and appears to be a reliable alternative, particularly in practical settings where the presence and type of data distribution shift is often unknown.

Haya Elayan, M. Sperrin, G. Martin et al. · 0 citations
Open access 2026

Evaluating Uncertainty Quantification in Clinical Machine Learning: Calibration, Robustness, and Decision Utility under Distribution Shift

A rigorous empirical framework is presented for comparing three uncertainty quantification approaches on two clinical prediction tasks, in-hospital mortality and 30-day readmission, using 74,829 ICU admissions from the MIMIC-IV database to support a more demanding evaluation standard for UQ in clinical machine learning.

Isaac Tosin Adisa, Francis Mawutor Amuyao, Ezekiel Olaoluwa Joaquim · 0 citations
Review Open access Sep 2026

Clinical Prediction Models in Cardiovascular Disease: Foundations, Clinical Applications and Future Directions.

Clinical prediction models have progressed from traditional risk scores derived from epidemiological cohorts into sophisticated systems capable of integrating multimodal data, electronic health records, and machine learning approaches. This review provides a clinically focused synthesis of the clinical prediction model lifecycle, outlining their conceptual foundations and methodological evolution across health care settings, with particular emphasis on cardiovascular risk prediction. These predictive frameworks inform both population-level public health strategies and individual clinical decision-making. However, external validation studies frequently reveal a stark discrepancy between statistical performance and real-world implementation, driven by temporal calibration drift, geographic transportability failures, and systemic sociodemographic or sex-based biases. Importantly, clinical prediction models should not be viewed as replacements for randomized trial evidence; instead, they function as vital complementary mechanisms to dissect the heterogeneity of treatment effects and individualize treatment decisions by contextualizing average treatment effects at the bedside. As health care shifts toward data-driven infrastructures, this review aims to contextualize the bench-to-bedside translational gap, offering an evidence-based roadmap for clinicians, statisticians, and data scientists to safely harness predictive modeling and ensure meaningful improvements in patient outcomes.

Bing-Yi Wang, A. Franciosi, L. O'Neill et al. · 0 citations
Open access Sep 2026

Evaluating discrimination of prediction models for recurrent clinical events

Prediction models for clinical conditions involving recurrent events require performance metrics tailored to repeated outcomes. Although we and others have developed methods to evaluate calibration and related aspects of performance, approaches for assessing discrimination remain limited. Recent contributions, including concordance-based methods and Brier-type accuracy measures, address parts of this gap, but no widely adopted, flexible framework for discrimination exists. As discrimination is a core TRIPOD+AI–recommended metric, accessible methodology and software for evaluating how well recurrent event models separate higher- and lower-risk individuals are still needed. We propose adapted discrimination statistics based on comparing predicted and observed cumulative event counts, addressing limitations of conventional concordance approaches. Statistical uncertainty is evaluated using several resampling strategies, including delete-d jackknifing, with user-friendly R code provided. We illustrate the approach using four tie-handling methods (c statistic, Kendall’s τₐ, Somers’ D, Goodman–Kruskal’s γ), four recurrent event models (negative binomial, zero-inflated negative binomial, Andersen–Gill, and Prentice–Williams–Peterson Total Time), and two clinical datasets: the SANAD epilepsy trial and an OPCRD asthma cohort. Discrimination varied substantially across models. In the asthma data, the Prentice–Williams–Peterson model showed the strongest discrimination (C = 0.939, 95% CI 0.935–0.944), with rank-based metrics (τₐ = 0.603; Somers’ D and γ = 0.879) also indicating very strong discrimination. The Andersen–Gill model performed moderately well (C = 0.823, 95% CI 0.815–0.830), while the negative binomial (C = 0.573) and zero-inflated negative binomial models (C = 0.586) showed weak discrimination. In the epilepsy data, the Prentice–Williams–Peterson model again performed best (C = 0.943, 95% CI 0.912–0.974), with all rank-based metrics above 0.85. The negative binomial, zero-inflated negative binomial and Andersen–Gill models showed moderate discrimination. This study provides practical methods and open-source code for evaluating the discrimination of prediction models for recurrent event data. Used alongside existing calibration tools, these methods support comprehensive performance evaluation and help standardise validation and reporting for prediction models developed for recurrent clinical events as recommended by the TRIPOD + AI guidelines.

T. Spain, Alexandra Hunt, Maria Sudell et al. · 0 citations
Preprint Aug 2026

From Test Performance to Risk-Based Effect Sizes: A Unified Wald-Type Framework to Design Clinical Validation Studies for Binary and Survival Outcomes

Clinical validation studies of predictive tests are usually designed to focus on sensitivity ($Se$) and specificity ($Sp$), while statistical power is often calculated on regression-effect scales (e.g., risk ratio, hazard ratio). However, these quantities are statistically connected. Here, we provide closed-form links from sensitivity, specificity, and disease prevalence ($\pi$) to predictive risks, risk contrasts, and Wald-type variance, power, and sample-size formulas for binary and fixed-horizon survival outcomes. Analyses of statistical efficiency via C- and D-optimal principles demonstrate how prevalence and threshold choices affect study efficiency, supporting rapid decisions in preliminary studies and informing the design of subsequent, larger studies. Simulations show good calibration across most realistic scenarios; when events are rare and test effects are simultaneously very large, continuity and minimum-event corrections are needed to stabilize the approximation. We illustrate the framework with a case study describing use of the coronary artery calcium score for predicting incident cardiovascular disease in patients with type 2 diabetes mellitus. The formulas let investigators check power and required enrollment directly from $(Se,Sp,\pi)$, without running a separate simulation for each design candidate.

Yong-Qi Zhong, A. Hartman, Jing Zhang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.