Skip to content

FAIR-SHEPHERD: Fairness beyond Statistical Parity toward Structural Alignment

Aug 2026 · ACM Journal on Responsible Computing · 0 citations · 9 references

TL;DR

This work introduces FAIR-SHEPHERD, a structural policy-based framework for transparent fairness in real-world settings with noisy labels and shifting distributions and develops S-agnostic tools that optimize policy-aligned gradient objectives using lattice-defined proxies.

Abstract

All fairness algorithms unavoidably rely on normative assumptions about fair treatment, yet these assumptions often remain implicit. We argue that these assumptions should be formalized as explicit, auditable policies and introduce FAIR-SHEPHERD, a structural policy-based framework for transparent fairness in real-world settings with noisy labels and shifting distributions. FAIR-SHEPHERD uses gradients as attribution signals, encoded in a Structural Fairness Specification (SFS) that defines vertical coherence and orthogonality to vulnerable proxies. We introduce SFS metrics: Vertical Coherence Score (VCS) to measure directional coherence across adjacent normative slices, and Horizontal Leakage Score (HLS), augmented by a signed directional variant, to detect gradient alignment with policy-declared vulnerable or proxy directions. To enforce these policies, we develop S-agnostic tools that optimize policy-aligned gradient objectives using lattice-defined proxies. We demonstrate that outcome-based auditing is brittle to measurement error: under 10% label noise on Adult, Worst-Group AUC for ERM drops by 0.159. In contrast, our gradient-based structural metrics provide a label-agnostic audit of the decision logic, remaining stable even when evaluation labels are corrupted. Specialized fairness baselines including ARL, JTT, and GoG retain substantial structural leakage on COMPAS, with HLS values from about 0.47 to 0.71. Some also reduce EOD relative to ERM, which shows that outcome and structural criteria can diverge. Using our S-agnostic Gradient Penalty tools, we reduce policy-specified structural leakage by roughly 85% on COMPAS (0.478 → 0.073) and over 94% on Adult (0.138 → 0.008) while retaining competitive AUC and Worst-Group AUC. We have released our source code to facilitate further research here.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

Fair Like Us? Auditing LLM Alignment in Resource Allocation

It is found that large language models tend to prefer stricter fairness constraints than humans, show more self-interested behavior, are sensitive to how information is framed, and are difficult to align with human judgments using fine-tuning with current datasets.

Qi-Shen Han, Hadi Hosseini, Joshua Kavner et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Position: Fairness Failure in Generative Models is an Evaluation Problem

This position paper argues that fairness failures in generative models, albeit driven by multiple factors, are ultimately stemming from an evaluation problem: fairness findings are rarely comparable across papers or actionable for deployment decisions.

M. Vladimirova, Jean-Yves Franceschi, Thibaut Issenhuth · 0 citations
Open access Aug 2026

Fairness Invariants: A Relational Approach to Explaining and Mitigating Fairness Bugs

REMI, a framework for the automated localization, explanation, and mitigation of individual discrimination, is presented, which introduces a bidirectional relational explanation framework that learns over paired examples $(x, x')$ to identify regions of the input space where fairness is violated.

Ranit Debnath Akash, Ashish Kumar, Gang Tan et al. · 0 citations
#machine learning Preprint Sep 2026

Efficient Fairness Auditing Across Guidance Scales in Text-to-Image Diffusion Models via Causal Abstraction

Fairness auditing of text-to-image diffusion models often requires generating large numbers of images across sampling configurations, making comprehensive evaluation computationally expensive. We propose a causal-abstraction-based audit instrument for efficiently evaluating fairness under interventions on the classifie...

Nabila Tasfiha Rahman, Rajatsubhra Chakraborty, De-Peng Xu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Efficient Active Auditing of Multi-Group Fairness with Bias Probes

Over the past decade, Machine Learning (ML) has been trained under dual objectives: minimizing prediction error via Empirical Risk Minimization (ERM) while controlling unfairness bias. In practice, however, fairness-aware training often yields limited improvements over standard ERM, making reliable post hoc auditing es...

Ayoub Ajarra, Debabrota Basu · 0 citations
#artificial intelligence Preprint Sep 2026

One Example Is Enough to Pass Fairness Benchmarks: Rethinking Fairness Evaluation for Aligned LLMs

It is argued that BBQ-style multiple-choice abstention benchmarks measure a single structural cue, and a model that solves them does not thereby become fair, and is called for evaluation suites that cover a broader spectrum of fairness alignment.

Nai-Hao Deng, Samee Arif, Shuai-Chen Chang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.