Who bears the cost of fairness? Empirical evidence from a cross-regional bias detection and mitigation pipeline in machine learning
Abstract
Introduction: Machine learning systems deployed across socioeconomically and geopolitically heterogeneous regions frequently generate unequal error distributions while still appearing compliant under aggregate evaluation metrics. This creates major governance concerns for high-stakes applications such as credit scoring, employment screening, and healthcare decision-making. Objective: This study evaluated whether commonly used bias mitigation strategies can reduce cross-regional algorithmic disparities without imposing disproportionate performance costs on structurally disadvantaged populations. Methods: A data-centric bias detection and mitigation pipeline was implemented using a synthetic cross-regional dataset (N = 6 000) parameterised from World Bank, ILO, and Global Findex distributional statistics across six world regions: Sub-Saharan Africa, South Asia, East Asia, Western Europe, North America, and Latin America. Four fairness metrics were evaluated simultaneously across all regions: AUC-ROC, Equalized Odds, Demographic Parity Gap, and F1 Score. Results: Pre-processing reweighting produced no measurable fairness improvement across regions. Post-processing isotonic calibration reduced disparity in some high-resource regions but degraded performance in South Asia. False Positive Rates exceeded 0,913 across all regions, revealing threshold collapse, meaning models classified nearly all cases as positive despite high aggregate recall. Fairness intervention costs were geographically asymmetric: Sub-Saharan Africa experienced an F1 reduction of 0,014 for only a 0,007 reduction in Demographic Parity Gap, whereas North America remained largely unaffected. Conclusions: Cross-regional algorithmic fairness cannot be achieved through post-hoc correction alone. AI governance frameworks should therefore require region-stratified auditing, geographically disaggregated reporting, and region-specific fairness evaluation prior to deployment.