Skip to content
Conference

Pattern Stability Analysis in Post-Crisis SME Lending: A Time-Split Classification Framework

Aug 2026 · 2026 International Conference on Smart Data, Intelligence, and Analytics (ICoSDIA) · pp. 1-6 · 0 citations · 34 references

Abstract

Concept drift is a continuing challenge in risk modelling for Small and Medium Enterprise (SME) loans, as default patterns change significantly during economic downturns. Most current studies use Random K-Fold validation, which leaks future data into the training and overestimates the performance metrics. To this end, we applied a Time-Split Classification Framework to assess the loan default risk pre- and post-2008 financial crisis in temporally reliable conditions. Our pipeline involved the collection of the U.S. SBA loan dataset (1987-2014), data cleaning, strict chronological splitting, class imbalance correction using SMOTE, model training, and feature importance tracking. We chose Light Gradient Boosting Machine (LightGBM) as the main classifier and benchmarked it against Random Forest and Logistic Regression. Results show that random validation inflates the AUC score by 1.17 percentage points (approximately $1.21 \%$), and LightGBM under time-split evaluation achieved an AUC of 0.9649. Pattern Stability Analysis showed that loan duration and borrower location were the top predictors in both periods, but after the crisis lenders put less emphasis on employee headcount and more on total loan size and projected job creation. These findings establish that developing reliable SME loan default prediction systems requires the validation of chronological models.

View source

Similar papers

Open access Sep 2026

Credit Default Prediction Using Large Language Models and Machine Learning: An Application to Colombia’s Solidarity Sector

Credit default prediction is a standard risk-management task, and large language models (LLMs) have been proposed as prompt-based alternatives, without task-specific parameter updating, for institutions that cannot deploy full machine learning (ML) pipelines. This study evaluates the Informed GPT on Colombian solidarit...

Javier André Ferro Pérez, M. Arias-Serna, J. Quiza-Montealegre · 0 citations
Open access Aug 2026

PREDICTIVE MACHINE LEARNING MODELS FOR MITIGATING NON-PERFORMING LOANS IN EMERGING BANKING SYSTEMS

This study benchmarks four supervised machine learning classifiers — logistic regression, random forest, gradient boosting, and extreme gradient boosting (XGBoost) — in predicting twelve-month-ahead loan delinquency using a loan-level panel drawn from commercial banks operating in an emerging Central Asian banking syst...

Djamalov Gofir Oribjanovich · 0 citations
Open access Sep 2026

An Efficient Loan Default Risk Assessment Framework Using TPE-Optimized CatBoost Model

Reliable estimation of loan default risk plays a vital role in ensuring financial system resilience and enabling sound credit decision-making in today’s lending landscape. Although conventional statistical techniques offer transparency, they frequently struggle to model the intricate nonlinear patterns embedded in high...

Baidyanath Sou · 0 citations
Conference Aug 2026

Machine learning-based risk prediction of corporate digital transformation

Digital transformation creates long-term opportunities for firms, but it may also generate short-term financial pressure during implementation. Existing studies mainly examine the ex-post effect of digital transformation on firm performance, while relatively few focus on Ex-Ante risk prediction. To address this gap, th...

Jinqi Xie · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.