Skip to content

Principled Bias Detection and Mitigation across the ML Data Lifecycle

Unknown authors
· 0 citations · 17 references

TL;DR

This dissertation treats data bias as a first-class data management problem and develops principled frameworks for detecting and correcting it across the ML data lifecycle, introducing Uniform Bias (UB), an interpretable intersectional measure with formal guarantees, and an ILP-based mitigation framework that models coverage and budget constraints.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Efficient Active Auditing of Multi-Group Fairness with Bias Probes

Over the past decade, Machine Learning (ML) has been trained under dual objectives: minimizing prediction error via Empirical Risk Minimization (ERM) while controlling unfairness bias. In practice, however, fairness-aware training often yields limited improvements over standard ERM, making reliable post hoc auditing es...

Ayoub Ajarra, Debabrota Basu · 0 citations
Open access Sep 2026

Software Fairness Analysis and Repair via Causal Model-Guided Data Mutation

This work introduces Fairabel, a novel approach that repairs fairness by mutating the training data under the guidance of causal models and multi-objective optimization, and shows that Fairabel reduces ML software bias by 50% on average across three fairness metrics, outperforming the state-of-the-art with a 9% relativ...

Ying Xiao, Zhen-Peng Chen, Yepang Liu et al. · 0 citations
Review Sep 2026

Mitigating Algorithmic Bias in AI-Driven Decision Systems for Financial Services

Artificial intelligence (AI) and machine learning (ML) systems are increasingly embedded in financial services, powering credit scoring, lending, fraud detection, and risk management. While these technologies enhance efficiency and predictive performance, they also introduce algorithmic bias that can undermine fairness...

P. Busch, Trang Tracy Nguyen · 1 citation
Preprint Sep 2026

Fair Variable Selection

Algorithms are increasingly being used to help automate and improve data-driven decisions, but care must be taken to prevent such algorithms from learning discriminatory patterns from historical data and perpetuating their biases. Statistical notions of fairness aim to mitigate either a model's disparate impact on disa...

D. Sulem, Jack Jewson · 0 citations
Preprint Aug 2026

Unmasking Removal-Budget Confounding: A Matched Operating-Point Evaluation Framework for Adaptive Data Cleaning

Experiments on CIFAR-10 and ImageNet-100 demonstrate that most performance differences observed in naive evaluation shrink or vanish at low-to-moderate corruption when operating points are matched, highlighting that adaptive cleaning methods must be benchmarked at matched operating points to ensure performance gains re...

Wei-Hsiang Chen, Pin-Hsuan Yu, Chenglin Fang et al. · 0 citations
Open access

Non-invasive fairness drift monitor for machine learning models using conformance constraints

Machine learning models deployed in consequential domains can become unfair toward protected subgroups as the data they receive drifts over time, yet the protected attributes needed to measure fairness directly are often unavailable at runtime due to privacy regulation and operational constraints. This creates a gap: e...

R. A. Bush · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.