Skip to content

FairFund-Bench: Evaluating Distributive Bias in LLM Resource Allocation

Jul 2026 · arXiv.org · Vol abs/2607.28934 · 1 citation · 42 references
Computer Science

TL;DR

FairFund-Bench is introduced, a benchmark that systematically varies key features of previous audit designs: the evaluation task (rating, ranking, or allocation), comparison context (single or multi-stimulus), and whether the audit is transparent or disguised, indicating that current LLMs robustly reproduce human deservingness evaluations.

Abstract

Large language models (LLMs) are increasingly involved in the distribution of scarce resources, raising concerns about biased allocations based on characteristics like race and gender. Recent LLM audits have produced inconsistent results, however, finding evidence of both positive and negative discrimination towards women and ethnic minorities, even for the same models. We show that this disagreement can arise from differences in audit format and introduce FairFund-Bench, a benchmark that systematically varies key features of previous audit designs: the evaluation task (rating, ranking, or allocation), comparison context (single or multi-stimulus), and whether the audit is transparent or disguised. The benchmark comprises 600 requests for financial assistance created from human-authored templates (calibrated against 1.3M real GoFundMe campaigns) across three domains, four race and two gender categories, and five causal framings of need derived from welfare deservingness theory. Across 14 models, audit format changes the direction of bias: models advantage minorities when rating claimants individually but penalize some groups when ranking them side by side. Bias magnitude, though small overall, is several times greater in disguised audits than in transparent ones, where, faced with appeals differing only in claimants'names, models overwhelmingly split funds equally. Causal framing effects, by contrast, exceed demographic effects by roughly an order of magnitude and are consistent across models and audit formats, indicating that current LLMs robustly reproduce human deservingness evaluations. The benchmark scores models on four criteria (demographic bias, deservingness alignment, cross-task consistency, and cross-context consistency), is publicly available, and can be readily adapted to other substantive domains.

View source

Similar papers

Conference Open access 2026

Bias and Fairness in LLM-Based Recruitment: A Systematic Review

A PRISMA 2020-guided systematic literature review draws on 82 studies selected from 493 records retrieved from Scopus and Web of Science and reveals a structural disconnect in the fairness-in-NLP and HCAI governance literature.

Asmae El Moutafail, Khalid Belkhoutout · 0 citations
Open access Jul 2026

WHEN FAIR AI BECOMES UNFAIR: A COUNTERFACTUAL AUDIT OF POSITIONAL BIAS IN LARGE LANGUAGE MODELS FOR HIRING DECISIONS

Findings indicate that state-of-the-art LLMs can achieve a high degree of demographic neutrality; fundamental artefacts such as positional bias can nonetheless produce severely discriminatory outcomes; and bias auditing must extend beyond demographic parity to interaction artefacts and ecosystem structure.

A. Camargo, Rafaela Silva Figueiredo Camargo · 0 citations
#large language models Review Open access Sep 2026

A review of bias detection and fairness auditing techniques in LLMs

A thorough literature review is provided to encapsulate prior research on bias identification and fairness auditing, categorizing the findings according to various stages of study and proposing a unified pipeline for dataset integration and a modular framework for bias auditing.

Nani Kartik Kaveti, T. Pattanshetti · 0 citations
#artificial intelligence Preprint Sep 2026

The Audit Decides the Verdict: Instrument Effects Rival Demographic Bias in LLM Decision Audits

Whether a language model looks demographically biased can depend on how the audit asks its question. A charitable-aid benchmark reports that the same models favor minority applicants when rating requests one at a time and penalize some when ranking side by side. We test whether that reversal generalizes to hiring, lend...

Siddharth Vohra, Manikandan Ravikiran · 0 citations
Preprint Aug 2026

Poli-Bias: Understanding and Measuring Large Language Model Biases in International Political Conflicts

Measuring political bias in large language models (LLMs) remains challenging as it can manifest through subtle differences in framing, argumentation, and legal reasoning that are difficult to capture with a single metric. In this work, we introduce Poli-Bias, a counterfactual framework for measuring whether LLMs treat...

Massi-Nissa Abboud, Aladin Djuhera, Elena Cabrio et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.