Churn Prediction and Risk Profiling Using Machine Learning and Customer Segmentation
Abstract
Customer churn occurs when they stop doing business with a company, leading to the loss of customers. Customer churn prediction has become a significant technique for identifying customers who are about to churn. The proposed method begins with data preprocessing, involving the standardization of data types and the imputation of missing values. Subsequently, exploratory data analysis is performed to allow us to have a better understanding of the data. This is followed by the use of the Synthetic Minority Oversampling Technique (SMOTE) to balance both the churner and non-churner classes. In this paper, the performance of three types of classifiers: logistic regression, random forest, and XGBoost Classifier are compared using the publicly available E-commerce and Bank Churners datasets. SHapley Additive exPlanations (SHAP) is utilized to identify the relevance of features and to demonstrate how features contribute to the model prediction. Next, K-means clustering is used to divide customers into distinct segments, and finally Bayesian logistic regression is applied to perform cluster risk analysis of each cluster, thus verifying and validating the risk profiles of them. Across all datasets, XGBoost delivers the best results, followed by random forest and logistic regression. The highest accuracy of 98.99% is achieved on the E-commerce dataset, while the Bank Churners dataset achieves a notably high accuracy of 98.36%. Besides, the results with and without SMOTE highlight the importance of balancing classes in getting better results. After using SMOTE, the F1 score and recall show marked improvement.