Sep 2026· IIARD INTERNATIONAL JOURNAL OF BANKING AND FINANCE RESEARCH· 0 citations
TL;DR
This study demonstrates that PLSDA-optimized machine learning achieves competitive accuracy with reduced computational complexity and enhanced interpretability in churn prediction, while meeting regulatory compliance requirements for practical banking implementations.
Abstract
Customer churn poses significant challenges to banking profitability, with acquiring new
customers costing 5-7 times more than retaining existing ones. Traditional statistical methods
struggle with high-dimensional data and complex non-linear patterns. This study addresses these
limitations by integrating advanced dimensionality reduction techniques with optimized machine
learning to achieve accurate, interpretable churn prediction. Three dimensionality reduction
techniques—UMAP, NCA, and PLSDA—were evaluated using a 10,000-customer European
banking dataset across ten machine learning classifiers: SVM, Logistic Regression, KNN,
Gaussian Naive Bayes, Decision Tree, Random Forest, AdaBoost, Bagging, Stacking, and
Voting. Data preprocessing included SMOTE balancing, StandardScaler normalization, and
one-hot encoding. GridSearchCV optimized hyper parameters systematically. Dual evaluation
employed 80-20 train-test split and 10-fold cross-validation. SHAP framework provided
comprehensive explainability through five visualization techniques. PLSDA emerged as the
superior approach, achieving 84.68% accuracy with Random Forest using 8 features—20%
dimensionality reduction while retaining 99% performance. Cross-validation confirmed
robustness: 85.27 ± 0.50% accuracy, 92.79 ± 0.70% ROC-AUC. Ensemble methods
outperformed single classifiers consistently. SHAP analysis identified Age, Number of Products,
and Balance as dominant churn drivers. This study demonstrates that PLSDA-optimized machine
learning achieves competitive accuracy with reduced computational complexity and enhanced
interpretability. The framework provides actionable insights for targeted retention strategies
while meeting regulatory compliance requirements for practical banking implementations.
Three dimensionality reduction techniques are employed to combine 10 machine learning classifiers including SVM, KNN, logistic regression, etc., to conduct research on the prediction of credit card customer churn to provide practical insights for financial institutions aiming to deploy efficient customer retention mod...
Benjamin Chiemeka Opara· International Journal of Eco...· 0 citations
This research has proposed a novel Hilbert-Schmidt Independence Criterion (HSIC) amidst other techniques for the selection of the intricate features for a robust predictive performance, allowing banks to better personalize service approaches to keep clients.
Benjamin Chiemeka Opara· IIARD International Journal...· 0 citations
This study compared explainable machine learning models for predicting customer churn using the IBM Telco Customer Churn dataset in R and found Logistic Regression achieved the best performance, with an accuracy of 82.30% on this dataset.
Uppu Venkata Subbarao, Tedlapu Narayana Rao, Vantaku Bala et al.· International Journal of Man...· 0 citations
Customer churn remains one of the most consequential problems facing business enterprises globally. Compared with the enticement and acquisition costs associated with acquiring new customers, retaining an existing subscriber is substantially cheaper. Due to its significance to business sustainability, various studies h...
Fatima Labake Ajani, O. A. Alimi, S. Moyane et al.· Informatics· 0 citations
This paper investigates the effectiveness of feature selection techniques in optimizing supervised
machine learning pipelines for customer churn prediction using the publicly available Customer
Churn Dataset from Kaggle. Feature selection plays a crucial role in enhancing model
interpretability and generalization by...
M. C. Opara· IIARD International Journal...· 0 citations
The Extreme Gradient Boosting algorithm is applied to telecom customer churn prediction, comparing its performance with Logistic Regression and Random Forest using the public Telco Customer Churn dataset and showing XGBoost outperformed benchmark models.
Hao Song· ITM Web of Conferences· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.