Skip to content
Book Open access

MAPPE: Rethinking and Improving Fairness in LLMs for Medicine via Minimax Preference-based Prompt Evolution

Aug 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · pp. 12832-12843 · 0 citations · 21 references

Abstract

Large Language Models (LLMs) have shown strong potential in medical applications such as question answering and clinical prediction. % Despite their growing adoption, fairness in LLMs for medicine remains underexplored, largely due to the mismatch between conventional fairness constraints and the clinically meaningful role of sensitive attributes. Existing approaches often enforce attribute-invariant constraints, leading to substantial performance degradation that is unacceptable in high-stakes healthcare settings. Moreover, fairness evaluation for medical LLMs is hindered by the lack of dedicated benchmarks. In this paper, we first rethink fairness in LLMs for medicine from a clinically grounded, utility-based perspective. Inspired by principles of health equity in medicine, we introduce universal fairness, a clinically grounded definition that reframes fairness as maximizing subgroup-aware diagnostic performance under attribute-conditioned health disparities. To achieve this objective in practice, we propose MAPPE, a training-free minimax prompt optimization framework. % MAPPE theoretically promotes universal fairness, while directly applicable to both closed-source and open-source LLMs. To systematically evaluate fairness, we construct FairMed, the first attribute-annotated benchmark for medical LLMs covering medical question answering and clinical prediction. % Experiments on both closed-source and open-source LLMs reveal demographic disparities, while MAPPE consistently improves worst-group and overall performance, outperforming existing fairness-oriented and prompt-based methods. The dataset and code are available at https://github.com/xiye7lai/FairMed.

Read PDF

Similar papers

Sep 2026

FairMoE-Health: Fairness-Aware Mixture of Experts for Equitable Multimodal Clinical Prediction.

Mixture-of-Experts (MoE) models for multi modal clinical prediction route patients to specialized ex pert networks based on input modalities, but we show that this routing mechanism introduces a previously un recognized source of demographic bias: because modality availability (e.g., whether a chest X-ray exists) correlates with race, gender, and insurance status, the gating network learns routing patterns that systematically differ across demographic groups. We propose FairMoE Health, a fairness-aware MoE framework that intervenes at three architectural levels: adversarial debiasing of encoded representations reduces demographic information before it reaches the gating network; adversarial debiasing of gating weights further discourages demographic leakage into routing decisions; and equalized odds regularization provides a prediction-level safety net. We introduce the Routing Disparity (RD) metric to quantify demographic imbalance in expert assignments. On three clinical prediction tasks from MIMIC-IV (mortality, length-of-stay, readmission), FairMoE-Health consistently reduces racial Equalized Odds Difference (EOD) and RD across all three tasks, with only small AUROC decreases. The Routing Disparity metric and multi-level debiasing framework introduced here generalize to MoE systems operating on demographically heterogeneous populations, providing both an audit tool and an architectural intervention for a bias mechanism that existing fairness methods leave unaddressed.

Xiaoyang Wang, Christopher C. Yang · 0 citations
Review Open access Aug 2026

Navigating fairness in artificial intelligence-based prediction models: theoretical constructs and practical applications.

Artificial intelligence (AI)-based prediction models, including risk scoring systems and decision support systems, are being increasingly adopted in health care. Addressing AI fairness is essential to fighting health disparities and ensuring equitable model performance and patient outcomes. However, numerous and conflicting definitions of fairness complicate this effort. In this Viewpoint, we aim to support the transition of AI fairness from theory to practice using appropriate fairness metrics. We assess the relation of 27 fairness definitions identified in the literature to the model's intended use, type of decision influenced, and ethical principles of distributive justice. Because of limitations in some notions of fairness, we argue that clinical utility, performance-based metrics (such as area under the receiver operating characteristic curve), calibration, and statistical parity are the most relevant group-based metrics for medical applications. Through two use cases, we show that different metrics might be applicable depending on the intended use and ethical framework. Our approach provides practical guidance for fair AI development, helping AI developers and assessors to evaluate model fairness and understand the effects of bias mitigation strategies, thereby supporting equitable AI-based implementations.

S. L. van der Meijden, Yuqing Wang, Madelena Y. Ng et al. · 1 citation
Open access Aug 2026

MaternaAI: Enhancing Equitable Maternal Healthcare in Kerala with Fairness-Aware and Explainable Learning Models

This paper introduces MaternaAI, a fairness-aware and explainable learning framework designed to enhance maternal healthcare predictions in Kerala, India, and proposes Adaptive Equity Score Optimization (AESO), a novel optimization algorithm that dynamically integrates fairness constraints into model training.

A. A. Jo · 0 citations
#software testing Open access Aug 2026

FIFT: Feature Importance-Guided Fairness Testing for machine learning software

Results indicate that global feature importance, used as an active search signal rather than a post-hoc diagnostic, improves both the effectiveness and the efficiency of individual fairness testing.

H. Mamman, Abdullateef Oluwagbemiga Balogun, Mustapha Maidawa et al. · 0 citations
#artificial intelligence Preprint Aug 2026

FairGlucose: A CGM Fairness Benchmark Reveals Subgroup Disparities Hidden in Population-Level Validation

It is established that population-level validation alone is insufficient for equity assessment of digital health AI, motivating subgroup-disaggregated reporting as a default standard, and subgroup-disaggregated reporting as a default standard for personalized configurations.

Junjie Luo, Xuzhe Zhi, Rui Han et al. · 0 citations
Review Open access Aug 2026

Challenges in medical algorithmic fairness

Abstract Artificial intelligence (AI) is increasingly being integrated into oncology for applications including cancer detection, risk stratification, treatment planning, and clinical documentation. Concerningly, growing evidence demonstrates that AI systems can reproduce or amplify existing disparities across patient populations. Although considerable effort has focused on developing computational methods to reduce algorithmic bias, many challenges surrounding fairness extend beyond technical implementation. In this commentary, we examine algorithmic fairness in oncology from technical and normative perspectives. We review common sources of bias throughout the machine-learning pipeline; discuss major statistical definitions of fairness, including demographic parity, calibration, and equalized odds; and highlight the inherent trade-offs among these metrics. We further explore how fairness often conflicts with overall predictive performance, arguing that model selection inevitably reflects ethical judgments rather than purely technical optimization. We discuss the limitations of current bias mitigation strategies and contend that many disparities rooted in historical and structural inequities cannot be resolved through algorithmic interventions alone. Finally, we outline priorities for the responsible development and deployment of clinical AI, including greater transparency in fairness decisions, context-specific evaluation standards, ongoing postdeployment auditing, and stronger regulatory oversight. Achieving equitable AI in oncology will require coordinated efforts among developers, clinicians, regulators, and patients to ensure that these technologies improve outcomes without perpetuating existing inequities.

Anand Srinivasan, D. Sritharan, S. Aneja et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.