Skip to content
Open access

Development of a machine learning model for loan default risk prediction in savings and credit cooperative organizations

Sep 2026 · African Journal of Empirical Research · 0 citations · 14 references

Abstract

Loan default remains a substantial challenge for Savings and Credit Cooperative Organizations (SACCOs), as it can increase credit risk, reduce profitability, and weaken financial sustainability. Although SACCOs increasingly use computerized information systems to manage lending activities, historical borrower data are often not fully utilized for predictive credit-risk assessment. This study developed an explainable machine learning model for loan default risk prediction in SACCOs. The study was guided by Information Asymmetry Theory and Credit Rationing Theory, which provide a basis for understanding how borrower information can reduce uncertainty and support credit allocation decisions. Four objectives guided the study: to identify factors associated with loan default risk, develop machine learning models for loan default prediction, evaluate their predictive performance, and provide an interpretable model for credit-risk assessment. A quantitative predictive analytics design was adopted using 1,000 records from the German Credit Dataset. Data preparation involved cleaning, categorical encoding, feature scaling, feature engineering, and class balancing using the Synthetic Minority Oversampling Technique (SMOTE). Logistic Regression and Random Forest models were developed using hyperparameter optimization and Stratified 5-Fold Cross-Validation. Model performance was evaluated using Accuracy, Precision, Recall, F1-score, ROC-AUC, Brier Score, confusion-matrix analysis, and false-negative analysis. SHapley Additive exPlanations (SHAP) was subsequently used to interpret the selected model. The results indicated that checking account status, age, credit amount, and loan duration were important predictors of default risk. Logistic Regression outperformed Random Forest across all reported evaluation measures, achieving Accuracy of 77.0%, Precision of 60.6%, Recall of 66.7%, F1-score of 63.5%, ROC-AUC of 0.795, and Brier Score of 0.162. It also recorded fewer false-negative predictions, with 20 compared with 26 for Random Forest. SHAP analysis identified checking account status and age as the most influential predictors. The study concludes that Logistic Regression with SHAP-based explainability provides a practical and interpretable approach to credit-risk prediction under the conditions examined. However, the model should be validated using institution-specific SACCO data before operational implementation.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.