2026· International Conference on Data Technologies and Applications· pp. 15-24· 0 citations· 13 references
Computer Science
TL;DR
This paper proposes a comprehensive Machine Learning pipeline that bridges the gap between predictive performance and model interpretability and integrates SHAP-based Explainable Artificial Intelligence to provide both global and local interpretability, revealing that contract type, tenure, and technical support subscriptions are the primary churn drivers.
Abstract
: Customer churn poses a significant threat to profitability in the saturated telecommunications industry, yet accurately predicting churn remains challenging due to high-dimensional data and class imbalance where churners represent a small minority. While complex ensemble methods achieve high accuracy, their ”black-box” nature limits business adoption, as practitioners require transparent insights to design effective retention campaigns. This paper proposes a comprehensive Machine Learning pipeline that bridges the gap between predictive performance and model interpretability. We implement and compare five classifiers—Logistic Regression, Random Forest, XGBoost, LightGBM, and CatBoost—optimized via Bayesian hyperparameter tuning and evaluated using the recall-focused F2-score to address class imbalance. Our results demonstrate that gradient boosting models, particularly XGBoost, outperform aggregation-based ensemble strategies, achieving the highest F2-score of 0.7500 with a recall of 0.9385. Crucially, we integrate SHAP-based Explainable Artificial Intelligence to provide both global and local interpretability, revealing that contract type, tenure, and technical support subscriptions are the primary churn drivers. This dual-layer transparency enables targeted retention strategies while preserving high recall in imbalanced telecom data.
The Extreme Gradient Boosting algorithm is applied to telecom customer churn prediction, comparing its performance with Logistic Regression and Random Forest using the public Telco Customer Churn dataset and showing XGBoost outperformed benchmark models.
These findings demonstrate that combining heterogeneous algorithms yields a reliable boost in predictive accuracy for both churn and potential return, informing more cost-effective retention and win-back strategies.
The results show that the proposed system can not only achieve stable predictive performance but also reveal the key drivers of customer churn, such as "engagement score" and "days since last purchase", which can provide actionable, data-driven business strategies for customer retention.
Yun-Hao Leng· Applied and Computational En...· 0 citations
This study compared explainable machine learning models for predicting customer churn using the IBM Telco Customer Churn dataset in R and found Logistic Regression achieved the best performance, with an accuracy of 82.30% on this dataset.
Uppu Venkata Subbarao, Tedlapu Narayana Rao, Vantaku Bala et al.· International Journal of Man...· 0 citations
A CRM system that relies on clever machine learning techniques and outperforms state-of-the-art machine learning and deep learning models for predicting customer attrition so that businesses may pinpoint consumers who are likely to churn, which allows for more proactive retention measures, happier customers, and more p...
Ramendra Pratap Singh· International Journal of Int...· 0 citations
Experimental results demonstrate that the proposed framework achieved strong and balanced predictive performance and demonstrated improvements in several classification metrics compared with the previously developed BFA-based Hybrid TabNet-DNN model.
Abdulrashid Abdulrauf, M. M. Lawal, Oluwatoyin Omoloba et al.· Scientific Journal of Comput...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.