Skip to content
#software testing Open access

Customer Churn Prediction Using Machine Learning Models: A Comparative Study Using R Software

Aug 2026 · International Journal of Management and Humanities · 0 citations · 11 references

TL;DR

This study compared explainable machine learning models for predicting customer churn using the IBM Telco Customer Churn dataset in R and found Logistic Regression achieved the best performance, with an accuracy of 82.30% on this dataset.

Abstract

Predicting customer churn is essential for improving retention and supporting long-term business growth. In this study, we compared explainable machine learning models for predicting customer churn using the IBM Telco Customer Churn dataset in R. Our approach included data preprocessing, exploratory analysis, model development, performance evaluation, and further analysis. We developed and evaluated four classification algorithms: Logistic Regression, Decision Tree, Random Forest, and Extreme Gradient Boosting, using an 80:20 train-test split. We assessed each model’s accuracy, precision, recall, and F1-score. Logistic Regression achieved the best performance, with an accuracy of 82.30% on this dataset. Feature importance analysis indicated that contract type, customer tenure, monthly charges, total charges, and internet service were the key factors influencing churn. These results suggest that explainable machine learning offers both strong predictive performance and greater transparency. The R-based framework we present provides a practical, reproducible approach to support customer retention strategies and help managers make evidence-based decisions in customer relationship management.

Read PDF

Similar papers

Open access Aug 2026

Optimizing bank customer churn prediction with machine learning models: a comparative study

This study utilizes five machine learning algorithms to classify and predict churn behavior based on financial and demographic characteristics from a dataset of 10,000 customers of a confidential multinational bank. This dataset is publicly accessible on Kaggle. The models in this study are random forest (RF), decision...

T. Tran, Thinh-Tien Bui · 0 citations
Open access Aug 2026

Predictive Analytics for Customer Retention: Designing and Evaluating a Machine Learning–Based Churn Prediction System

Results demonstrate that ensemble models, particularly those trained using the Random Forest and Gradient Boosting algorithms, outperform baseline approaches across all selected evaluation metrics, and these algorithms are recommended for identifying potential churners across various business and industrial use cases.

Blerina Çeliku, Marijon Pano · 0 citations
Open access Sep 2026

Comparative Performance of Supervised Machine Learning Models for Customer Churn Prediction Using E-Commerce Behavioral Data

Customer churn remains a major challenge for e-commerce organizations because customer retention directly affects profitability and business sustainability. This study presents a comparative evaluation of five supervised machine learning algorithms for customer churn prediction using 3,000 customer records obtained fro...

Ahmed Sani Yeldu, Rufai Aliyu Yauri, Sirajo Abdullahi Bakura et al. · 0 citations
Open access Sep 2026

Reducing Model Complexity in Bank Customer Churn Prediction Using Dimensionality Reduction and Explainable Machine Learning

This study demonstrates that PLSDA-optimized machine learning achieves competitive accuracy with reduced computational complexity and enhanced interpretability in churn prediction, while meeting regulatory compliance requirements for practical banking implementations.

Prisca Chimezie Opara · 0 citations
Open access Sep 2026

Machine Learning–Based Customer Churn Prediction in Banking Using Feature Selection and Ensemble Models

This research has proposed a novel Hilbert-Schmidt Independence Criterion (HSIC) amidst other techniques for the selection of the intricate features for a robust predictive performance, allowing banks to better personalize service approaches to keep clients.

Benjamin Chiemeka Opara · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 2, 2026

Documenting the tech worker movement

Writing as a participant and researcher, PhD student JS Tan SM ’22 has co-authored a new book about the rise of tech worker protests and the employer backlash that followed.

GPT-Lab Sep 23, 2026

Requirements Don’t Live in Isolation: What We’re Exploring with Req-Space

Requirements in large systems rarely exist in isolation. Their meaning depends on the wider project context - other requirements, policies, decisions, tests, and implementation details. That becomes especially important when AI is used for review, because spotting a possible conflict or gap is only the beginning. ReqSpace explores how AI, visualisation, and connected project context can help reviewers understand those findings, trace the relationships behind them, and focus on the questions that…

GPT-Lab Sep 17, 2026

Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering

AI is making software generation faster, but speed does not remove the need for expertise. As more work is delegated to AI, tacit knowledge may become one of the most important human advantages in software engineering. The post Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering appeared first on GPT-Lab.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.