Skip to content
Preprint

Beyond Aggregate Calibration: Decomposing Income-Conditional Recall Disparities in Automated Credit Default Prediction

Aug 2026 · 0 citations · 6 references
Computer Science

TL;DR

Empirical findings show that simply blinding an algorithm to sensitive attributes fails to ensure fairness when institutional pricing decisions and behavioral proxy variables collectively reconstruct the omitted signals, and outline the practical implications for auditing data-centric AI workflows within regulated financial institutions.

Abstract

Data-centric curation pipelines frequently rely on model confidence scores to flag and filter noisy or mislabeled training instances. Evaluating this filtering convention on a large-scale consumer lending sample (LendingClub, N = 1,344,936) uncovers an underlying demographic asymmetry: high-income defaulters are disproportionately classified as label noise relative to low-income defaulters (Cramer's V approximately 0.03-0.07). Re-examining this behavior through the lens of equal opportunity [Hardt et al., 2016] reveals a far more severe discrepancy: a 16.86 percentage point gap in true positive rate (recall) between high- and low-income borrowers who ultimately defaulted. Implementing a sequential feature-blinding methodology allows us to isolate the drivers of this disparity across three distinct mechanisms: (1) direct reliance on self-reported applicant income; (2) algorithmic absorption of upstream institutional bias encoded within origination interest rates; and (3) a residual disparity (3.55 percentage points in cross-validation; 2.56 percentage points on a held-out test partition, Z = -4.04, p<0.0001) that remains even after purging both income and interest rates from the model. Out-of-sample signed SHAP valuations demonstrate that this residual gap is maintained by structural proxies, most notably loan amount and home ownership status. These empirical findings show that simply blinding an algorithm to sensitive attributes fails to ensure fairness when institutional pricing decisions and behavioral proxy variables collectively reconstruct the omitted signals. We outline the practical implications of these findings for auditing data-centric AI workflows within regulated financial institutions.

View source

Similar papers

Open access Jul 2026

UPI behavioral proxies for thin-file credit risk prediction

India's Unified Payments Interface (UPI) generates rich behavioral data for over 400 million active users, yet this transactional signal remains inaccessible to most lenders for credit assessment due to data privacy restrictions. Thin-file borrowers — individuals with limited formal credit history — represent the primary beneficiaries of UPI-based credit assessment and simultaneously the population for whom bureau-based models perform least reliably. This study proposes, empirically validates, and evaluates six UPI behavioral proxy features (F4) constructed from standard loan application variables, providing a publicly replicable framework for approximating UPI transaction signals in credit scoring. Scientific validation via Spearman correlation confirms that payment_discipline_score and upi_success_ratio_proxy exhibit the expected directional relationships with default in both independent datasets (p<0.001). Using DS1: LendingClub 2016-2018 (N=300,001) and DS2: Home Credit (N=307,511) with five machine learning models, results show F4 features consistently improve default detection Recall for thin-file borrowers across all models and both datasets. CatBoost achieves +9.95% Recall gain (DS1) and Logistic Regression +6.80% (DS2). SHAP attribution confirms F4 proxies account for 19.0%–32.2% of total predictive power for thin-file borrowers, with payment_discipline_score ranking as the single most predictive feature in Home Credit above all bureau variables. These findings establish UPI behavioral proxies as a meaningful, scientifically validated, and previously underquantified dimension of creditworthiness.

Deep Shikha, Himanshu Vasnani · 0 citations
Review Open access Jul 2026

CreditR1: Calibration-Aware Reinforcement Learning for Interpretable Corporate Credit Risk Assessment with Large Language Models

CreditR1 delivers calibrated PDs with evidence-grounded reasoning that supports internal model validation and human review that supports transferability beyond the Chinese A-share market remains an open empirical question.

Yuxuan Wu, Haowen Dai, Yiheng Zhang et al. · 0 citations
Preprint Aug 2026

Interpretable hybrid credit scoring for thin-file and underbanked populations

We extend a residual-learning hybrid credit scoring framework (logistic regression scorecard plus a gradient-boosting correction on its residuals, decomposed at each prediction into an interpretability ratio $\rho(x)$ that measures the share attributable to the linear branch) along three axes: an East African empirical instantiation on the Zindi Financial Inclusion in Africa data (Kenya, Rwanda, Tanzania, Uganda); a fairness audit at the granularity of the framework's three interpretability regions; and a thin-file segmentation analysis. On the Taiwan Credit Default benchmark retained for continuity, the calibrated hybrid attains AUC $= 0.776$ ($\Delta\mathrm{AUC} = +0.057$ vs.\ standalone logistic regression, $+0.001$ vs.\ standalone XGBoost), reduces Brier Score by 23\%, and concentrates the highest-default-rate borrowers (69.5\%) in the fully interpretable region. On Zindi, the calibrated hybrid attains AUC $= 0.869$ ($\Delta\mathrm{AUC} = +0.015$ vs.\ LR, $p<0.001$; $-0.004$ vs.\ XGBoost), cuts Brier from $0.158$ to $0.085$ (a 46\% reduction), and replicates the regional routing pattern. The fairness audit detects severe routing into the opaque ML-driven region along socioeconomic axes: rural respondents by 18 percentage points relative to urban, primary-or-less-educated by 32 points relative to secondary-and-above, and Ugandan respondents by 22 points relative to Kenyan, while gender shows essentially no routing disparity. The audit pipeline surfaces subgroup-routing violations that aggregate fairness metrics miss, in a form directly usable by African central-bank supervisors of digital credit.

Belise Kanziga, Yaé U. Gaba, Olivier Kanamugire · 0 citations
Conference Open access Jul 2026

Selection Bias Correction in Retail Intelligence

Retail intelligence often relies on monitoring popular, high-velocity products, potentially biasing economic indicators by ignoring the"long tail"of niche items. This simulation study investigates selection bias in inflation estimation and compares correction methods across diverse data-generating processes. Through 400 Monte Carlo replications spanning four scenarios--aligned step functions, smooth gradients, misaligned breaks, and polynomial relationships--we test the robustness of Inverse Probability Weighting (IPW) with five specifications against stratification with varying strata counts. Our findings reveal fundamental limits of weighting methods in retail long-tail contexts: stratification achieves superior performance in three of four scenarios, maintaining sub-0.04pp median error even when boundaries deliberately misalign with population breaks (116x advantage over IPW). However, IPW with spline propensity models wins under smooth polynomial relationships (median error 0.007pp vs. 0.013pp), demonstrating context-dependency. Critically, even an oracle IPW specification with perfect structural knowledge achieves 6.06pp error compared to stratification's 0.008pp in step-function scenarios. This reflects violation of the Positivity Assumption--a fundamental causal inference requirement--rather than IPW methodological inferiority. When selection probabilities differ dramatically (90% vs. 1%), weighting methods operate outside their theoretical design envelope. These results demonstrate that stratification provides a safer engineering choice in retail long-tail distributions with severe positivity violations.

S. Chowdhury · 0 citations
Conference Open access 2026

Behavioral biases and artificial intelligence in banking decision-making: Toward explainable hybrid systems for SME financing

SME credit files arrive incomplete, and the gaps leave room for anchoring, confirmation bias and loss aversion. We compare human, algorithmic and hybrid credit decisions using a benchmark credit dataset alongside a vignette experiment with credit analysts working in Morocco's Souss-Massa region. The modelling arm pairs L2-regularised logistic regression with gradient-boosted trees, adding stratified validation, calibration analysis, SHAP and LIME. In the human arm, matched cases vary the requested amount while everything else is held constant. Analysts were least stable on borderline files, and their decisions moved with the anchor. The boosted model held steadier but leaned harder on indicators that track how thick a file is. AI-first assistance improved consistency and deepened deference to the model; human-first assistance preserved contextual overrides; explanation-gating struck the best balance, though only where SHAP and LIME agreed. We assess distribution through demographic-parity difference, disparate-impact ratio, equal-opportunity difference and false-positive-rate difference. What the results support is a governed hybrid: weak explanations withheld, overrides auditable, human review genuinely available. A regional sample and benchmark data bound how far any of these travels.

Hassan Ennaqui, Mohamed El Bourki, Abdellah Bakrim et al. · 0 citations