Skip to content
Open access

From Fairness Findings to Fairness Claims: An Evidence Classification Scheme for Clinical AI

Jul 2026 · medRxiv · 0 citations
Medicine

TL;DR

An evidence classification scheme is introduced that screens for sample size and precision, and integrates stability across design alternatives directly into the fairness claim, on the estimation of the brain-age gap from structural MRI using the Alzheimer's Disease Neuroimaging Initiative data.

Abstract

Fairness audits of clinical AI models rarely make the evidentiary status of subgroup findings explicit: reassuring results may reflect insufficient statistical precision rather than true parity, and audit verdicts can easily reverse under equally defensible analytic choices. We introduce an evidence classification scheme that screens for sample size and precision, and integrates stability across design alternatives directly into the fairness claim. We demonstrate this scheme on the estimation of the brain-age gap (BAG), a potential clinical biomarker, from structural MRI using the Alzheimer's Disease Neuroimaging Initiative (ADNI) data. The male-female and Black-vs-White differences, along with the White-Male and Black-Female intersectional contrasts, are all classified as equivalence supported, stable across regressor choice (ridge vs. gradient-boosted trees) and feature representation (full feature set vs. cortical-thickness-only). The Asian-vs-White and Black-Male comparisons remain classified as insufficient data throughout, as neither meets the pre-specified minimum-sample threshold. The proposed scheme provides a path from raw fairness findings to justified fairness claims via pre-specified thresholds, minimum-information screening, and stability checks across declared design choices.

Read PDF

Similar papers

#artificial intelligence Preprint Aug 2026

FairGlucose: A CGM Fairness Benchmark Reveals Subgroup Disparities Hidden in Population-Level Validation

It is established that population-level validation alone is insufficient for equity assessment of digital health AI, motivating subgroup-disaggregated reporting as a default standard, and subgroup-disaggregated reporting as a default standard for personalized configurations.

Junjie Luo, Xuzhe Zhi, Rui Han et al. · 0 citations
Open access Aug 2026

A Simplified Metric to Streamline Between-Group Fairness Assessment for Predictive Models: Algorithm Development and Evaluation Study

Abstract Background Fairness evaluation is essential for trustworthy clinical risk prediction. However, existing fairness-oriented discrimination metrics either ignore cross-group comparisons or rely on exhaustive pairwise evaluations, making them difficult to interpret and impractical for model selection. Objective Th...

Hao-Yuan Wang, Chuan Hong, Michael J. Pencina et al. · 0 citations
Open access Aug 2026

Fairness Evaluation Paradox: How Biased Test Data Masks True Group Fairness Assessment

The EU AI Act makes fairness metrics for high-risk AI systems’ compliance evidence, turning their trustworthiness into a safety and accountability concern. Fairness audits assume that test data reflects the properties of real-world conditions, whereas standard evaluation protocols use test data drawn from the same bias...

Sašo Karakatič, Ivona Colakovic, Tjaša Heričko · 0 citations
Preprint Sep 2026

Instability Floors: Separating Bias from Noise in Fairness Audits of Clinical LLM Agents with FairMedAgent

Counterfactual fairness audits of clinical language-model agents report a flip rate: how often an action changes when only the patient's demographic descriptor changes. Part of that rate is not demographic. A stochastic agent also changes its own action when nothing changes, and a flip rate cannot be interpreted withou...

Rohith Reddy Bellibatlu, Manpreet Singh, Deepak Parashar et al. · 0 citations
#machine learning Preprint Oct 2026

Beyond Demographic Balance: Multi-Metric and Intersectional Evaluation of Fairness in MIMIC-IV Mortality Prediction

Fairness conclusions in clinical prediction can depend strongly on both the metrics reported and the demographic resolution at which performance is evaluated. We revisit these evaluation choices for ICU mortality prediction on MIMIC-IV, comparing predictive-utility and subgroup-error metrics across several fairness int...

A. Al Noman, Fahmid Al Rifat, Tahrima Hashem et al. · 0 citations
Open access Sep 2026

Designing fairness: best practices for gender-sensitive development of cognitive ability tests in recruitment

Psychological tests play a pivotal role in high-stakes decisions such as recruitment, yet traditional development guidelines concentrate fairness work downstream—at test administration and post hoc statistical bias correction—while offering little concrete guidance for the design stage, where constructs are defined and...

Jannick Schneider, Clemens Striebing, Melanie Elizabeth Jacobsen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.