Skip to content
Preprint

Fairness-Aware Test-Time Prompt Tuning

Aug 2026 · 0 citations · 51 references
Computer Science

TL;DR

FairTPT is developed, a novel fairness-aware episodic TTA method that jointly minimizes target marginal entropy while maximizing spurious marginal entropy through soft-prompt tuning and establishes a foundation for robust TTA, which is essential for achieving fairness in practice.

Abstract

Vision-language models have displayed remarkable capabilities in multi-modal understanding and are increasingly used in critical applications where economic and practical deployment constraints prohibit re-training or fine-tuning. However, these models can also exhibit systematic biases that disproportionately affect protected demographic groups and existing approaches to addressing these biases require extensive model retraining and access to demographic attributes. There is a clear need to develop test-time adaptation (TTA) approaches that improve the fairness characteristics of pretrained models under distributional shift. In this paper, we evaluate how episodic TTA affects fairness in CLIP classification under subpopulation shifts and develop FairTPT, a novel fairness-aware episodic TTA method that jointly minimizes target marginal entropy while maximizing spurious marginal entropy through soft-prompt tuning. We find that standard episodic TTA generally exacerbates disparities between majority and minority groups, that blinding a model to spurious attributes without degrading target performance is inherently challenging, and that excessive blinding can lead to catastrophic forgetting. This model collapse can be prevented by monitoring test-time changes in target loss within the linear regime, while still achieving fairness improvements on reactive data and preserving overall performance. FairTPT outperforms all state-of-the-art episodic test-time debiasing methods and establishes a foundation for robust TTA, which is essential for achieving fairness in practice.

View source

Similar papers

Preprint Aug 2026

Fairness-Aware Mixture-of-Experts via Subgroup Reweighting and Gate Regularization

This work identifies routing-induced bias, a failure mode in which subgroup imbalance drives the gating network to route subgroups onto a few experts, and proposes an end-to-end Mixture-of-Experts (MoE) framework that corrects it and improves fairness while maintaining competitive predictive performance.

Sunhee Hwang · 0 citations
Book Open access Sep 2026

From Constraint to Control: Modeling Expected Fairness in Ranking Systems

Modern learning to rank systems have achieved remarkable performance across a wide range of applications. However, they may also exhibit disparities in exposure, raising concerns about fairness, especially in sensitive domains such as healthcare, judicial decision-making, and recruitment, where biased rankings may have...

Tristan Cladière, Antoine Gourru, Bissan Audeh et al. · 0 citations
Preprint Aug 2026

Follow the Norm: Accounting for Fine-Tuning and Prompt Effects on Model Rationales

It is demonstrated in controlled experiments that norm-breaking fine-tuning yields norm-divergent actions justified by self-interested rationales, suggesting a systematic shift in patterns of justification.

Long Hoang Nguyen, Brice Valentin Kok-Shun, Guangyu Du et al. · 0 citations
Open access Sep 2026

Software Fairness Analysis and Repair via Causal Model-Guided Data Mutation

This work introduces Fairabel, a novel approach that repairs fairness by mutating the training data under the guidance of causal models and multi-objective optimization, and shows that Fairabel reduces ML software bias by 50% on average across three fairness metrics, outperforming the state-of-the-art with a 9% relativ...

Ying Xiao, Zhen-Peng Chen, Yepang Liu et al. · 0 citations
Preprint Aug 2026

Training Fair Tabular Foundation Models

Tabular Foundation Models (TFMs) have emerged as leading methods for tabular predictive tasks, leveraging in-context learning to predict on new data without task-specific training. Despite the increased use of TFMs in high-stakes decision-making, their fairness properties remain largely unexplored. In this work, we inc...

P. Kenfack, Jesse C. Cresswell, Anthony L. Caterini et al. · 0 citations
Preprint Aug 2026

Taming the Implicit: Dual-Channel Risk-Aware Reinforcement Fine-Tuning for Continual Multimodal Post-Training

Reinforcement fine-tuning (RFT) is widely believed to inherently resist catastrophic forgetting in continual post-training of multimodal large language models. Under pronounced task distributional shifts, however, forgetting across representative RFT algorithms escalates sharply. This stems from the implicit reward-var...

Yibei Liu, Jiajun Chen, Qianle Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.