Skip to content
Conference

Churn Prediction and Risk Profiling Using Machine Learning and Customer Segmentation

Jul 2026 · International Conference on Information and Communicatiaon Technology · pp. 1-6 · 0 citations · 22 references

Abstract

Customer churn occurs when they stop doing business with a company, leading to the loss of customers. Customer churn prediction has become a significant technique for identifying customers who are about to churn. The proposed method begins with data preprocessing, involving the standardization of data types and the imputation of missing values. Subsequently, exploratory data analysis is performed to allow us to have a better understanding of the data. This is followed by the use of the Synthetic Minority Oversampling Technique (SMOTE) to balance both the churner and non-churner classes. In this paper, the performance of three types of classifiers: logistic regression, random forest, and XGBoost Classifier are compared using the publicly available E-commerce and Bank Churners datasets. SHapley Additive exPlanations (SHAP) is utilized to identify the relevance of features and to demonstrate how features contribute to the model prediction. Next, K-means clustering is used to divide customers into distinct segments, and finally Bayesian logistic regression is applied to perform cluster risk analysis of each cluster, thus verifying and validating the risk profiles of them. Across all datasets, XGBoost delivers the best results, followed by random forest and logistic regression. The highest accuracy of 98.99% is achieved on the E-commerce dataset, while the Bank Churners dataset achieves a notably high accuracy of 98.36%. Besides, the results with and without SMOTE highlight the importance of balancing classes in getting better results. After using SMOTE, the F1 score and recall show marked improvement.

View source

Similar papers

Conference Open access 2026

Customer Churn Prediction in Telecommunications Using Machine Learning Models

. Customer churn prediction has a great impact on the long-term profitability and competitiveness of telecommunications companies. This study uses telecommunications customer data, employs machine learning methods to build a customer churn prediction model, and identifies key influencing factors. By using majority clas...

Jian-Qing Cao · 0 citations
Open access Sep 2026

Machine Learning–Based Customer Churn Prediction in Banking Using Feature Selection and Ensemble Models

This research has proposed a novel Hilbert-Schmidt Independence Criterion (HSIC) amidst other techniques for the selection of the intricate features for a robust predictive performance, allowing banks to better personalize service approaches to keep clients.

Benjamin Chiemeka Opara · 0 citations
#software testing Open access Aug 2026

Customer Churn Prediction Using Machine Learning Models: A Comparative Study Using R Software

This study compared explainable machine learning models for predicting customer churn using the IBM Telco Customer Churn dataset in R and found Logistic Regression achieved the best performance, with an accuracy of 82.30% on this dataset.

Uppu Venkata Subbarao, Tedlapu Narayana Rao, Vantaku Bala et al. · 0 citations
Open access Sep 2026

Performance Evaluation of Single and Ensemble Models for Customer Churn Prediction Analysis

Customer churn remains one of the most consequential problems facing business enterprises globally. Compared with the enticement and acquisition costs associated with acquiring new customers, retaining an existing subscriber is substantially cheaper. Due to its significance to business sustainability, various studies h...

Fatima Labake Ajani, O. A. Alimi, S. Moyane et al. · 0 citations
Conference Open access 2026

XGBoost Algorithm for Telecom Customer Churn Prediction and Its Business Implications

The Extreme Gradient Boosting algorithm is applied to telecom customer churn prediction, comparing its performance with Logistic Regression and Random Forest using the public Telco Customer Churn dataset and showing XGBoost outperformed benchmark models.

Hao Song · 0 citations
Open access Sep 2026

Reducing Model Complexity in Bank Customer Churn Prediction Using Dimensionality Reduction and Explainable Machine Learning

This study demonstrates that PLSDA-optimized machine learning achieves competitive accuracy with reduced computational complexity and enhanced interpretability in churn prediction, while meeting regulatory compliance requirements for practical banking implementations.

Prisca Chimezie Opara · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.