Explainable Machine Learning for Home Equity Line of Credit Risk Assessment: A Multi-Model Comparison with SHAP Interpretability
Abstract
Accurate default risk prediction is essential for commercial banks' retail credit loan decisions and supervision. Traditional credit scoring models fail to provide clear decision schemes under strict regulation, while high-performance machine learning models face the black-box dilemma, making it difficult to meet financial regulatory requirements for model transparency. Based on the HELOC benchmark data released by FICO, this study investigates machine learning applications in credit risk prediction and benchmarks six models: random forest, XGBoost, LightGBM, CatBoost, logistic regression, and SVM. Of the six models evaluated, CatBoost recorded the highest AUC at 0.7923; robustness assessments, however, showed that the margins separating the models fell short of statistical significance. We then applied SHAP to interrogate the basis of these predictions. Three features emerged as the dominant drivers of HELOC risk: external risk estimation, revolving credit burden, and the average length of credit history. For credit risk practitioners, these variables deserve the closest attention when reviewing HELOC applications. More broadly, the findings indicate that interpretable machine learning can give lenders a defensible account of every credit decision — a capability that serves both regulatory oversight and the individuals whose applications are under review.