Current churn prediction research mainly emphasizes static accuracy on curated datasets, and they often overlook label noise arising from heuristic CRM rules, uncontrollable user attrition, and suboptimal retention timing. This study proposes a robust two-stage retention framework targeting e-commerce users with tenure of more than four years. Firstly, structural churn cases are filtered from the dataset. Then, 5% label noise is added to simulate real-world annotation errors, and benchmark evaluations are conducted across nine machine learning models. Ensemble methods, specifically GBM, XGBoost, and LightGBM, achieve F1 scores ranging from 0.43 to 0.44 with cross-validation standard deviations below 0.02. About the method, the analysis progresses from binary classification to Cox proportional hazards modeling, identifying a proactive intervention window between the fourth and fifth year of user tenure. Business-oriented threshold tuning enables GBM campaigns to attain a 345.7% return on investment. This significantly surpasses simple methods. Further analysis via SHAP values shows service call fatigue is the primary churn driver. The assumed impact of fee sensitivity lacks empirical support within this mature cohort. The proposed project effectively bridges statistical prediction and operational retention, providing an effective methodology for protecting high-value customer assets.
Research Article
Open Access