Aug 2026· Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2· pp. 12832-12843· 0 citations· 21 references
Abstract
Large Language Models (LLMs) have shown strong potential in medical applications such as question answering and clinical prediction. % Despite their growing adoption, fairness in LLMs for medicine remains underexplored, largely due to the mismatch between conventional fairness constraints and the clinically meaningful role of sensitive attributes. Existing approaches often enforce attribute-invariant constraints, leading to substantial performance degradation that is unacceptable in high-stakes healthcare settings. Moreover, fairness evaluation for medical LLMs is hindered by the lack of dedicated benchmarks. In this paper, we first rethink fairness in LLMs for medicine from a clinically grounded, utility-based perspective. Inspired by principles of health equity in medicine, we introduce universal fairness, a clinically grounded definition that reframes fairness as maximizing subgroup-aware diagnostic performance under attribute-conditioned health disparities. To achieve this objective in practice, we propose MAPPE, a training-free minimax prompt optimization framework. % MAPPE theoretically promotes universal fairness, while directly applicable to both closed-source and open-source LLMs. To systematically evaluate fairness, we construct FairMed, the first attribute-annotated benchmark for medical LLMs covering medical question answering and clinical prediction. % Experiments on both closed-source and open-source LLMs reveal demographic disparities, while MAPPE consistently improves worst-group and overall performance, outperforming existing fairness-oriented and prompt-based methods. The dataset and code are available at https://github.com/xiye7lai/FairMed.
Mixture-of-Experts (MoE) models for multi modal clinical prediction route patients to specialized ex pert networks based on input modalities, but we show that this routing mechanism introduces a previously un recognized source of demographic bias: because modality availability (e.g., whether a chest X-ray exists) correlates with race, gender, and insurance status, the gating network learns routing patterns that systematically differ across demographic groups. We propose FairMoE Health, a fairness-aware MoE framework that intervenes at three architectural levels: adversarial debiasing of encoded representations reduces demographic information before it reaches the gating network; adversarial debiasing of gating weights further discourages demographic leakage into routing decisions; and equalized odds regularization provides a prediction-level safety net. We introduce the Routing Disparity (RD) metric to quantify demographic imbalance in expert assignments. On three clinical prediction tasks from MIMIC-IV (mortality, length-of-stay, readmission), FairMoE-Health consistently reduces racial Equalized Odds Difference (EOD) and RD across all three tasks, with only small AUROC decreases. The Routing Disparity metric and multi-level debiasing framework introduced here generalize to MoE systems operating on demographically heterogeneous populations, providing both an audit tool and an architectural intervention for a bias mechanism that existing fairness methods leave unaddressed.
Xiaoyang Wang, Christopher C. Yang· IEEE journal of biomedical a...· 0 citations
Artificial intelligence (AI)-based prediction models, including risk scoring systems and decision support systems, are being increasingly adopted in health care. Addressing AI fairness is essential to fighting health disparities and ensuring equitable model performance and patient outcomes. However, numerous and conflicting definitions of fairness complicate this effort. In this Viewpoint, we aim to support the transition of AI fairness from theory to practice using appropriate fairness metrics. We assess the relation of 27 fairness definitions identified in the literature to the model's intended use, type of decision influenced, and ethical principles of distributive justice. Because of limitations in some notions of fairness, we argue that clinical utility, performance-based metrics (such as area under the receiver operating characteristic curve), calibration, and statistical parity are the most relevant group-based metrics for medical applications. Through two use cases, we show that different metrics might be applicable depending on the intended use and ethical framework. Our approach provides practical guidance for fair AI development, helping AI developers and assessors to evaluate model fairness and understand the effects of bias mitigation strategies, thereby supporting equitable AI-based implementations.
S. L. van der Meijden, Yuqing Wang, Madelena Y. Ng et al.· The Lancet Digital Health· 1 citation
This paper introduces MaternaAI, a fairness-aware and explainable learning framework designed to enhance maternal healthcare predictions in Kerala, India, and proposes Adaptive Equity Score Optimization (AESO), a novel optimization algorithm that dynamically integrates fairness constraints into model training.
Results indicate that global feature importance, used as an active search signal rather than a post-hoc diagnostic, improves both the effectiveness and the efficiency of individual fairness testing.
H. Mamman, Abdullateef Oluwagbemiga Balogun, Mustapha Maidawa et al.· Journal of King Saud Univers...· 0 citations
It is established that population-level validation alone is insufficient for equity assessment of digital health AI, motivating subgroup-disaggregated reporting as a default standard, and subgroup-disaggregated reporting as a default standard for personalized configurations.
Junjie Luo, Xuzhe Zhi, Rui Han et al.· 0 citations
Abstract Artificial intelligence (AI) is increasingly being integrated into oncology for applications including cancer detection, risk stratification, treatment planning, and clinical documentation. Concerningly, growing evidence demonstrates that AI systems can reproduce or amplify existing disparities across patient populations. Although considerable effort has focused on developing computational methods to reduce algorithmic bias, many challenges surrounding fairness extend beyond technical implementation. In this commentary, we examine algorithmic fairness in oncology from technical and normative perspectives. We review common sources of bias throughout the machine-learning pipeline; discuss major statistical definitions of fairness, including demographic parity, calibration, and equalized odds; and highlight the inherent trade-offs among these metrics. We further explore how fairness often conflicts with overall predictive performance, arguing that model selection inevitably reflects ethical judgments rather than purely technical optimization. We discuss the limitations of current bias mitigation strategies and contend that many disparities rooted in historical and structural inequities cannot be resolved through algorithmic interventions alone. Finally, we outline priorities for the responsible development and deployment of clinical AI, including greater transparency in fairness decisions, context-specific evaluation standards, ongoing postdeployment auditing, and stronger regulatory oversight. Achieving equitable AI in oncology will require coordinated efforts among developers, clinicians, regulators, and patients to ensure that these technologies improve outcomes without perpetuating existing inequities.
Anand Srinivasan, D. Sritharan, S. Aneja et al.· JNCI Cancer Spectrum· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.