Skip to content
Conference

Mitigating Spurious Correlations in Text Classification Using Latent Space Geometry

2026 · Annual Meeting of the Association for Computational Linguistics · pp. 27525-27539 · 0 citations · 25 references
Computer Science

TL;DR

This paper introduces a prototype-guided modeling approach that leverages natural language prompts to represent confounders, transforming abstract biases into interpretable geometric anchors without auxiliary classifiers, a novel framework that mitigates spurious correlations by manipulating latent space geometry.

Abstract

Spurious correlations cause deep learning models to rely on predictive shortcuts that hold in the training data but break under distribution shifts, leading to large performance drops for minority groups. Existing strategies often rely on costly group annotations or employ unstable adversarial training. In this paper, we pro-pose Prototype-guided debiasing using Robust Invariant Feature Transformations (PRIFT), a novel framework that mitigates spurious correlations by manipulating latent space geometry. Specifically, we introduce a prototype-guided modeling approach that leverages natural language prompts to represent confounders, transforming abstract biases into interpretable geometric anchors without auxiliary classifiers. Based on these anchors, we introduce a centered projection operator that adaptively puri-fies representations by removing confounding deviations specific to instances while preserving essential semantic structure. Furthermore, PRIFT can handle confounding factor information at different levels, ranging from true labels to unsupervised latent inference. Experiments on four text classification benchmarks demonstrate the superiority of our method; notably, PRIFT outperforms state-of-the-art baselines and improves worst-group accuracy by over 20% on the CivilComments dataset compared to standard empirical risk minimization.

View source

Similar papers

#machine learning Preprint Sep 2026

SAGE: Subpopulation-Aware Generative Enhancement for Mitigating Spurious Correlations

Subpopulation-Aware Generative Enhancement (SAGE), a two-stage generative augmentation framework, is introduced, using cluster-derived sub-labels and class labels to fine-tune a conditional generative model and text encoder, generating targeted synthetic data to fill underrepresented regions in the training set and con...

Yi-Ming Luo, Rong-Qiang Zhao, Jie Liu · 0 citations
Preprint Aug 2026

UNMASK: Discovering and Causally Verifying Spurious Shortcuts in Text Classifiers

U N M ASK is presented, a fully automated pipeline that discovers, causally verifies, and mitigates spurious correlations in text classifiers without additional human annotation, and demonstrates that the discovery and validation stages generalize to reward model preference data.

Chidaksh Ravuru, Shashank Srivastava · 0 citations
2026

A Multi-Dimensional Evaluation of Explainability in Media Bias Detection

It is suggested that predictive performance, attribution plausibility, and mechanistic faithfulness characterize different aspects of model behavior and should be evaluated separately when studying explainability in media bias detection.

Tinghao Chen, Raina Zhang, Benjamin M. Ampel et al. · 0 citations
Conference 2026

ProReGen: Progressive Residual Generation under Attribute Correlations

ProReGen is presented, a progressive residual generation approach inspired by the classical Robinson’s transformation, to partial out from an image attribute x2 its component mx1 that is predictable by other image attributes x1, and the residual γ=x2-mx1 that is not.

Ruby Shrestha, Ajay Gopi, Casey Meisenzahl et al. · 0 citations
#machine learning Preprint Aug 2026

MERIT: Mitigating Exposure Bias in Generative XMC for User-Interest Propensity Modeling

Matching users to interest categories at scale is central to personalized shopping, but the task is challenging in large e-commerce platforms, where label spaces continually evolve and user-interest signals are sparse and long-tailed. Autoregressive language models are appealing because their world knowledge and semant...

Abhinav Mahajan, Arindam Sarkar, P. Comar · 0 citations
Jul 2026

D3O: Dynamic Distribution Distillation for Ordinal Regression

The proposed D3O, a dynamic distribution distillation framework that replaces static supervision with training-driven evolution of ordinal label distributions via self-distillation, introduces a contrastive ordinal-aware label enhancement module that leverages vision-language alignment to recover refined label distribu...

Chunlai Dong, Yao-Jun Hu, Yuyang Xu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.