Skip to content
Open access

Explainable Deep Tabular Learning for Credit Risk Assessment: An Information-Theoretic Cross-Attentional Transformer Approach

Jul 2026 · Entropy · Vol 28 · 0 citations · 19 references
Medicine

TL;DR

This study develops an explainable machine learning framework for modeling loan approval decisions on heterogeneous tabular data, centered on a Cross-Attentional Tabular Transformer that applies bidirectional cross-attention between numerical and categorical feature groups.

Abstract

Credit risk assessment is a core component of financial decision-making. This study develops an explainable machine learning framework for modeling loan approval decisions on heterogeneous tabular data, centered on a Cross-Attentional Tabular Transformer that applies bidirectional cross-attention between numerical and categorical feature groups. The prediction target is historical loan-approval status, treated as a proxy for, not a direct measure of, borrower default risk; a supplementary validation on a dataset with an authentic default label is also reported. Class imbalance is addressed through focal loss, and post hoc interpretability is provided through SHAP analysis. Three classifiers, Random Forest, Gradient Boosting, and the proposed transformer, are evaluated on a 5000-sample credit dataset using accuracy, precision, recall, F1-score, ROC-AUC, and average precision. Gradient Boosting achieves the best performance (accuracy 0.9640, F1-score 0.9189), with Random Forest comparable; the proposed transformer reaches 0.9530 accuracy and 0.8949 F1, without surpassing the ensembles and at substantially higher computational cost. A five-split robustness comparison additionally evaluates XGBoost, LightGBM, CatBoost, and calibrated logistic regression: all three Gradient-Boosting variants and both classical ensembles exceed the transformer’s performance on every metric, while calibrated logistic regression does not. The evaluated baseline set excludes deep tabular architectures such as TabNet, FT-Transformer, SAINT, and TabPFN-style methods. Across the three primary classifiers, SHAP identifies credit score, employment status, and income as the dominant features, consistent with domain expectations. The results characterize the observed performance–efficiency trade-off between ensemble methods and attention-based tabular learning under the evaluated data conditions.

Read PDF

Similar papers

Open access Aug 2026

SMOTETomek-DNN: A Machine Learning Framework for Credit Risk Prediction with an Imbalanced Dataset

The experimental findings demonstrate that incorporating resampling techniques substantially improved the default detection performance and suggest that hybrid sampling integrated with advanced learning architectures can provide a reliable and practical solution for managing credit risk in imbalanced microfinance datas...

Tiruneh Kebede Dubale, Siraj Sebhatu Seyoum · 0 citations
Open access Aug 2026

A TRANSFORMER-BASED APPROACH FOR FINANCIAL RISK ASSESSMENT

Financial risk assessment identifies and measures the risk of loss in financial decisions. It includes credit, market, cash flow, and operational risk. The traditional methods depend on history and expert judgment. Machine learning speed up and standardize risk prediction.  This study used Transformer model to classify...

Anam Naz, Shazia Batool, A. Qadoos et al. · 0 citations
Aug 2026

Explainable Ensemble Deep Learning for Credit Risk Classification

Experimental results on three real credit risk datasets show that the EED-CRC approach achieves superior performance compared to traditional CRC methods, both in accuracy and in explainability.

Sirine Ben Ghozzi, M-A. Ben Hajkacem, Nadia Essoussi · 0 citations
Open access 2026

CSA-Net: A Type-Aware Dual-Tower Neural Model for Imbalanced Tabular Credit Default Prediction

Credit default prediction under class imbalance requires attention to both ranking quality and fixed-threshold behavior. This study evaluates imbalanced tabular credit default prediction on the Default of Credit Card Clients (DCCC, approximately $4{:}1$ ) and Give Me Some Credit (GMSC, approximately $14{:}1$ ) benchmar...

Tharin Chab, Jaegi Jeon · 0 citations
Open access Aug 2026

Explainable Ensemble Learning for Loan Approval Prediction Using XGBoost, LightGBM, and Random Forest with SHAP Analysis

Loan approval prediction is a critical task in financial institutions, as it directly impacts risk management and decision-making processes. However, challenges such as class imbalance and lack of model interpretability often limit the effectiveness and reliability of machine learning approaches. This study proposes an...

Mikaria Gultom · 0 citations
Review Open access Sep 2026

Explainable Machine Learning for Home Equity Line of Credit Risk Assessment: A Multi-Model Comparison with SHAP Interpretability

Accurate default risk prediction is essential for commercial banks' retail credit loan decisions and supervision. Traditional credit scoring models fail to provide clear decision schemes under strict regulation, while high-performance machine learning models face the black-box dilemma, making it difficult to meet finan...

Kelvin Lin · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.