Jul 2026· International journal of computer information systems and industrial management applications· 0 citations
TL;DR
An Explainable Machine Learning (XML) framework for credit risk assessment that combines an ensemble classifier, integrating XGBoost, Random Forest, and LightGBM, with an integrated SHAP-and-LIME explainability layer is proposed and evaluated using a large-scale retail and priority-sector loan dataset drawn from public sector, private sector, regional rural, and small finance bank segments operating in India.
Abstract
Credit risk assessment remains a foundational function of commercial banking, yet the Indian banking sector's persistent non-performing asset (NPA) burden, heterogeneous borrower base, and expanding priority-sector and microfinance lending create distinct challenges that generic, globally trained credit scoring models do not adequately address. This paper proposes an Explainable Machine Learning (XML) framework for credit risk assessment that combines an ensemble classifier, integrating XGBoost, Random Forest, and LightGBM, with an integrated SHAP-and-LIME explainability layer, and evaluates the framework using a large-scale retail and priority-sector loan dataset drawn from public sector, private sector, regional rural, and small finance bank segments operating in India. The framework incorporates SMOTE-ENN based class-imbalance handling to address the low base default rate typical of retail lending portfolios, and an explainability-guided feature refinement step that uses SHAP attributions to iteratively prune low-value features and retain regulator-interpretable risk factors. The proposed model was evaluated on a dataset of 42,860 loan accounts and benchmarked against five baseline models. Results show that the proposed explainable ensemble achieves 93.2% accuracy, 91.0% precision, 89.4% recall, and a 90.2% F1-score, exceeding the strongest baseline (LightGBM) by 5.6 percentage points in F1-score, with an area under the ROC curve (AUC) of 0.967. Segment-wise analysis reveals materially higher default risk concentration in regional rural banks and small finance institutions relative to public and private sector banks, with debt-to-income ratio, credit bureau (CIBIL) score, and repayment delinquency history emerging as the most influential predictors across segments.
Credit default prediction has become an important application of machine learning in the banking and financial sector, as it helps financial institutions identify potential loan defaulters and support informed lending decisions. Although machine learning models often provide high predictive performance, many of them function as black-box system, making it difficult for financial analysts and decision-makers to understand the reasoning behind their predictions. This lack of transparency can reduce user trust, particularly in high-stakes financial applications where explainable decisions are essential. To address this challenge, this study explores the use of Explainable Artificial intelligence (XAI) techniques to improve the interpretability of credit default prediction. A Random Forest classifier was developed using a publicly available credit default dataset containing financial attributes such as employment status, bank balance, annual salary, and loan default status. The dataset was preprocessed and partitioned into training and testing sets before model development. To explain the prediction process, SHAP (Shapley Additive Explanations) and LIME (Local Interpretable Model-Agnostic Explanations) were integrated with the trained Random Forest model. SHAP was used to provide both global and local explanations by identifying the overall importance and contribution of individual features, while LIME generated instance-level explanations to illustrate how specific features influenced individual predictions. The explanation results were presented through visualizations, including feature importance plots, waterfall plots, and local explanation graphs, allowing a clearer understanding of the model's decision-making process. The findings demonstrate that the combined use of SHAP and LIME enhances the transparency and interpretability of the Random Forest model by providing complementary perspectives on feature contributions. This study highlights the practical value of explainable machine learning in developing more understandable, trustworthy, and accountable credit risk assessment systems for real-world financial decision-making.
Muskan, B. Sidhu· International Journal of Com...· 0 citations
It is argued that predictive accuracy and regulatory transparency are not competing objectives but complementary necessities for institutional survival in Nepal’s cooperative sector.
S. K. Sahani, Tsair-Fwu Lee, Digvijay Pandey et al.· Journal of Intelligent Decis...· 0 citations
The credit risk assessment remains a critical challenge for financial institutions as manual and semi-automated loan evaluation processes often produce inconsistent decisions, high default rates, and operational inefficiencies. This paper discusses an Integrated Data-Driven Loan Management Framework that comprises an Artificial Neural Network (ANN) for credit default prediction, Structured Query Language (SQL) to systematically extract and transform data, and Interactive Visual Analytics Dashboards to aid in providing transparency within the decision-making process. For the study, a public loan dataset was used to pre-process the target dataset containing a total of 38,577 records and 24 attributes with pre-processing techniques such as feature engineering, one-hot encoding and class re-balancing with the use of Synthetic Minority Oversampling Technique (SMOTE) yielding 43 model-ready inputs.
The final ANN architecture was developed with 2 hidden layers (43 and 21 neurons with ReLU activation) and a sigmoid output neuron for binary classification. SHapley Additive exPlanations (SHAP) were computed to provide interpretability for each prediction made by the model. The overall accuracy of the model is 87% and the weighted precision, recall, and F1-score are 0.87 for the held-out test set (n = 12,858). The framework is implemented as a web-based decision support tool using Flask that enables users to receive risk scores in real-time along with explainable outputs and visual dashboards. The results from this experimentation indicate that an integrated pipeline (including data querying, predictive modeling, interpretability, and visualization) provides better decision-making and stakeholder transparency than using a model in isolation; therefore, it serves as a proof-of-concept prototype for intelligent loan management in banks and other financial institutions.
Ponsak .S. Bande, Chimuzuoroke E. Ugwuja, Blessing .C. Uzo et al.· International Journal of Lat...· 0 citations
By empirically proving that high-performance algorithms can be mathematically blind to demographic biases, this framework directly advances SDG 10 (Reduced Inequalities) and provides the accountable, feature-level justifications required for secure and sustainable financial inclusion (SDG 8).
Htet Nge Nge Ko, Aung Htoo Khine, Shadab Kalhoro et al.· Journal of Risk and Financia...· 0 citations
Credit risk assessment forms a cornerstone of banking risk management and the stability of the wider financial system. Over the past decade, the rapid development of machine learning (ML) techniques has substantially enhanced traditional credit risk assessment methodologies. ML has now emerged as a core technological pillar for the banking sector, strengthening risk identification capabilities, optimising credit decision-making, and advancing financial inclusion. Conventional credit scoring models, dominated by logistic regression (LR) and scorecard approaches, offer inherent strengths in interpretability and regulatory compliance. However, constrained by their linear assumptions, these methods struggle to capture complex non-linear relationships within credit data and deliver insufficient predictive accuracy for the “credit-invisible” population lacking formal credit histories. This paper presents a systematic literature review (SLR) of ML applications in credit risk assessment (CRA), covering publications from January 2016 to May 2026. A total of 894 papers were retrieved from five digital libraries, and following a rigorous multi-stage screening process, 129 studies were selected for final inclusion. Our analysis reveals that tree-based ensemble models and deep learning (DL) architectures predominate in contemporary research in this field. Meanwhile, post hoc explanation methods and machine learning operations (MLOps) are gaining significant traction as solutions to address fairness, transparency, and system maintenance challenges in real-world production environments. We synthesise prevailing methodologies into a unified end-to-end credit risk modelling framework spanning data preprocessing, feature engineering, model training, evaluation, and operational deployment. Through a critical assessment of the advantages, limitations, and inherent trade-offs of existing approaches, this SLR not only identifies current research gaps and future directions for the academic community, but also provides practical guidance for the banking sector to build compliant, fair, and efficient intelligent risk assessment systems.
Bolun Zhang, Jun Luo, Ruobing Wu et al.· Journal of Risk and Financia...· 0 citations
Credit risk assessment is pivotal to the sustainability of micro-lending institutions, particularly in emerging economies such as Ghana, where conventional evaluation methods remain predominantly manual and subjective. Traditional approaches, which rely on face-to-face interviews, personal judgments, and simple background checks, are vulnerable to human biases, inconsistencies, and inefficiencies that contribute to elevated default rates and broader financial instability. This study investigates the application of machine learning (ML) techniques, specifically Random Forest (RF), Extra Tree Classifier (ETC), and a probability-averaged Ensemble Classifier, to enhance credit risk assessment in Ghanaian micro-lending institutions. Using a quantitative experimental research design, the study analysed 32,581 loan records drawn from Tepa Man Microfinance Institution. Data preprocessing included missing-value imputation, one-hot encoding, and class balancing via random oversampling, applied exclusively to the training set. Model performance was evaluated through 10-fold stratified cross-validation using accuracy, precision, recall, F1-score, AUC-ROC, Cohen's Kappa, and Matthews Correlation Coefficient (MCC). Hyperparameters were set to scikit-learn defaults (n_estimators = 100, random_state = 42) to ensure reproducibility. The Random Forest and Extra Tree Classifiers each achieved a mean accuracy of 99.33% and an AUC-ROC of 0.9997, results that are consistent with the high-quality, real-world dataset and are critically interpreted in the context of potential overfitting risks. Feature importance analysis identified the loan-to-income ratio and interest rate as the dominant predictors of default. The Ensemble Method, which averages class probabilities across both base models, achieved 84.25% accuracy and an AUC of 0.9231, demonstrating stronger generalization than the individual classifiers. The study concludes that integrating ML models can substantially improve the accuracy, consistency, and reliability of credit risk evaluations, thereby reducing default rates and supporting financial inclusion in Ghana's microfinance sector.
P. Addo, Samuel Kofi Akpatsa, Emmanuel Mensah et al.· Journal of Applied Social Sc...· 0 citations