Skip to content
Open access

The Fairness Illusion? A Cross-Dataset Audit of Accuracy and Demographic Bias in Credit Scoring Based on Machine Learning

Sep 2026 · Journal of Risk and Financial Management · Vol 19, pp. 674 · 0 citations · 32 references

TL;DR

Persistent demographic parity disparity among the more complex models is consistent with feature-level bias that no model architecture can resolve, and has direct implications for the less-discriminatory-alternatives framework under US fair lending law and for the high-risk classification of credit scoring AI under the EU AI Act.

Abstract

Machine learning has transformed consumer credit scoring, delivering substantial gains in predictive accuracy over traditional scorecards—but whether those gains come at a cost to fairness has remained contested. The dominant assumption in the literature is that more complex, accurate models amplify bias by encoding historical patterns of disadvantage more effectively. This paper challenges that assumption with direct empirical evidence. We evaluate four model families—logistic regression, random forest, XGBoost, and a multilayer perceptron—across two real-world datasets: the UCI Credit Card Default Dataset and the 2024 US Home Mortgage Disclosure Act national loan-level data, comprising over six million mortgage applications. Using repeated cross-validation, we report predictive performance alongside two primary fairness metrics—demographic parity difference and equalized odds difference—supplemented by false positive rate difference and calibration difference, with confidence intervals across 15 estimation folds. On the UCI data, where demographic disparities are modest, model choice has negligible effect on fairness outcomes. On the HMDA mortgage data, where racial disparities are large and legally consequential, the expected accuracy–fairness tradeoff does not hold; more accurate models produce significantly fairer outcomes on equalized odds within the models, data, and fairness criteria examined here, with logistic regression occupying the worst position simultaneously on all dimensions. Persistent demographic parity disparity among the more complex models is consistent with feature-level bias that no model architecture can resolve. The findings have direct implications for the less-discriminatory-alternatives framework under US fair lending law and for the high-risk classification of credit scoring AI under the EU AI Act.

Read PDF

Similar papers

Open access Oct 2026

Who bears the cost of fairness? Empirical evidence from a cross-regional bias detection and mitigation pipeline in machine learning

Introduction: Machine learning systems deployed across socioeconomically and geopolitically heterogeneous regions frequently generate unequal error distributions while still appearing compliant under aggregate evaluation metrics. This creates major governance concerns for high-stakes applications such as credit scoring...

Linda Bessa-Simons, Clinton Amponsah, B. Kyiewu · 0 citations
Open access Sep 2026

The Explainability–Reliability Gap in Fraud Detection: Evidence from SHAP and Permutation Importance Under Distribution Shift

This study develops an empirical audit framework for assessing the explainability–reliability gap in fraud detection: whether stable model explanations remain consistent with performance-based feature reliance under distribution shift. Using the Bank Account Fraud dataset suite, including a Base dataset and five biased...

Istiaque Bhuiyan, Rahmad Mirza, A. Hoque et al. · 0 citations
Review Open access Sep 2026

Algorithm Bias in Credit Markets and Socio-Economic Inequality – A Survey of The Six Geopolitical Zones in Niger

The rapid diffusion of machine-learning credit-scoring models promises faster loan decisions and lower transaction costs, yet mounting evidence suggests that such algorithms can reproduce and even exacerbate existing socio-economic disparities (Hand, 2021). In Nigeria, where formal financial inclusion varies markedl...

J. O. Onoh · 0 citations
#machine learning Preprint Sep 2026

A Rank Graduation metric for Algorithmic fairness

Fairness assessment in algorithmic decisions that affect individuals, such as credit scoring, often relies on parity measures calculated at the aggregate group level. Such measures may not reveal which individuals experience unfairness or which explanatory factors contribute to it. In this paper, we propose a rank-base...

Dalia Atif, Paolo Giudici · 0 citations
Review Open access Sep 2026

Explainable Machine Learning for Home Equity Line of Credit Risk Assessment: A Multi-Model Comparison with SHAP Interpretability

The findings indicate that interpretable machine learning can give lenders a defensible account of every credit decision — a capability that serves both regulatory oversight and the individuals whose applications are under review.

Kelvin Lin · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.