Algorithmic Fairness as a Risk-Management Problem in Banking and Insurance: Regulatory Frameworks, Model Governance, and Fairness-Aware Credit Scoring
Abstract
AI-driven credit scoring is supervised as a high-risk application in banking and insurance, yet unfairness is rarely operationalized as a measurable category of model, conduct, legal, and reputational risk. Using 20,000 anonymized applications from a Southern European digital lender (15.2% twelve-month default rate), we estimate three model families—a regularized logistic regression, a gradient-boosting machine, and a multi-layer perceptron—under a fully crossed design in which each family is evaluated without mitigation and under pre-processing (reweighing), in-processing (an exponentiated-gradient reduction, applicable to any base learner, together with adversarial debiasing where gradient-based training permits it), and post-processing (reject-option) interventions, so that the mitigation effect is no longer confounded with the choice of estimator. No sensitive-group field enters any estimated specification; group membership is used exclusively for auditing. Predictive performance (AUC-ROC, Brier score and Brier skill score relative to the base-rate forecast, F1 on the default class, Gini, and the Kolmogorov–Smirnov statistic) is reported jointly with group fairness (demographic-parity and equal-opportunity differences, disparate-impact ratio, Theil index) and with group-conditional calibration, at an explicitly stated and economically justified decision threshold. Every fairness quantity is accompanied by stratified-bootstrap confidence intervals and, for stochastic learners, by seed-level dispersion. The interpretable benchmark attains an AUC of 0.780 and a Brier score of 0.104 against 0.129 for the constant base-rate forecast, and the high-capacity models improve on it by under one AUC point. Disparity is present but is located geographically rather than in the composite group label: the disparate-impact ratio is 0.724 [0.693, 0.754] for the lowest socio-economic neighborhood cluster, excluding the four-fifths screening value, against 0.809 [0.776, 0.840] for the ethno-socioeconomic proxy, whose interval contains it, and no measurable gender disparity. Group membership is recoverable from the neutral feature set at an AUC of 0.654, and 42% of the group gap in predicted risk travels through the bureau credit score alone, so feature deletion cannot close the channel. Feature attributions and an auxiliary group-recoverability test locate the proxy pathways through which disparity arises, and a misclassification-sensitivity analysis bounds the effect of error in the group proxy, which attenuates measured disparity toward parity. We map the results onto Regulation (EU) 2024/1689 as amended by Regulation (EU) 2026/1744, the GDPR as interpreted in SCHUFA Holding, Directive (EU) 2023/2225, EBA loan-origination guidance, and Solvency II, EIOPA, and IAIS expectations, and propose fairness-risk controls organized around impact assessment, independent validation, and three lines of defense governance. Because the evidence comes from credit origination at a single lender, the insurance argument is developed at the level of regulatory and governance architecture rather than as an empirical transfer of estimates.