Jul 2026· Far East Journal of Electronics and Communications· Vol 30, pp. 153-174· 0 citations
TL;DR
This study addresses the critical trade-off between predictive accuracy and algorithmic fairness and proposes a “Three-Stage Fairness-Aware Framework” integrating pre-processing, in-processing, and post-processing mitigation strategies that successfully reduced bias to ethical thresholds.
Abstract
The integration of Artificial Intelligence (AI) into high-stakes social work, such as child welfare and criminal justice, promises efficiency but risks perpetuating systemic biases. This study addresses the critical trade-off between predictive accuracy and algorithmic fairness. Baseline evaluations of Extreme Gradient Boosting (XGBoost) models across three datasets (COMPAS, AFST, and a synthetic dataset) revealed strong predictive performance (AUC up to 0.89) but significant racial bias. As shown in the fairness metrics comparison, pre-mitigation Statistical Parity Difference (SPD) values indicated severe bias (COMPAS: –0.21, AFST: –0.15, Synthetic: –0.18). To resolve this, we propose a “Three-Stage Fairness-Aware Framework” integrating pre-processing, in-processing, and post-processing mitigation strategies. The application of this framework successfully reduced bias to ethical thresholds, with post-mitigation SPD values improving significantly to –0.07, –0.04 and –0.06, respectively, while incurring minimal accuracy loss (< 5.1% AUC reduction). Furthermore, the study validates the framework’s computational scalability and its alignment with strict regulatory standards like the EU AI Act and GDPR. These findings provide an evidence-based blueprint for equitable and legally compliant AI governance in public services.
The empirical findings support the use of complementary structural and policy-level interventions and demonstrate the importance of jointly evaluating aggregate disparity, worst-case attribute-level harm, cross-attribute transfer, and predictive utility.
It is argued that clinical utility, performance-based metrics, calibration, and statistical parity are the most relevant group-based metrics for medical applications and that different metrics might be applicable depending on the intended use and ethical framework.
S. L. van der Meijden, Yuqing Wang, Madelena Y. Ng et al.· The Lancet Digital Health· 1 citation
It is shown that racial bias is not only present in the COMPAS dataset but is also amplified by the models trained on it, which undercuts both the assumption that algorithmic decision-making offers a neutral improvement over human judgment and the weaker claim that it merely mirrors preexisting human bias.
Ignacio Cofone, Warut Khern-am-nuai University of Oxford, M. University· 3 citations
Persistent demographic parity disparity among the more complex models is consistent with feature-level bias that no model architecture can resolve, and has direct implications for the less-discriminatory-alternatives framework under US fair lending law and for the high-risk classification of credit scoring AI under the...
Colin Ellis· Journal of Risk and Financia...· 0 citations
AI-driven credit scoring is supervised as a high-risk application in banking and insurance, yet unfairness is rarely operationalized as a measurable category of model, conduct, legal, and reputational risk. Using 20,000 anonymized applications from a Southern European digital lender (15.2% twelve-month default rate), w...
Workforce evaluation is where organizational bias does its quietest damage: subjective peer reviews and managerial discretion systematically under-reward some groups and overlook “silent performers,” and the decisions are rarely explained. This paper presents FAIRLENS, a fairness-aware and explainable platform that aud...
V. S, Praveen G. R., William Arul Christofer V et al.· International Conference Com...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.