Back to projects

Customer Churn Prediction & Retention Optimization in Online Food Delivery

Analyze customer behavior and retention patterns in an online food delivery dataset, identify key factors associated with customer churn, and develop a machine learning solution to help prioritize customers for retention. The project combines exploratory data analysis, business insight generation, feature selection, predictive modeling, and interactive deployment.

Role

Data Analyst & Data Scientist

Timeline

2 Days

Tools

Python - Pandas - NumPy - Matplotlib - Seaborn - Scikit-learn - Jupyter Notebook - Streamlit - Joblib

Metabase dashboard placeholder.

This area is prepared for an embedded dashboard from Metabase. Later, replace the placeholder with the iframe link from your real dashboard.

MetabaseEmbed preview

Business context.

OrderKu needed to understand which customers were loyal and which customers were at risk of churn. With 388 customer records and a 22.4% overall churn rate, the analysis focused on uncovering behavioral and demographic patterns associated with churn, identifying high-risk customer segments, and translating the findings into actionable retention priorities.

Analysis approach.

  • Data Cleaning & Validation — cleaned categorical values, handled missing required fields, and checked duplicate records.
  • Exploratory Data Analysis — analyzed churn distribution across customer frequency, feedback, demographics, occupation, income, and location.
  • Customer Segmentation — compared Loyal vs At-Risk customers and identified customer personas based on behavioral and demographic patterns.
  • Feature Selection — excluded variables with potential leakage or limited modeling relevance, such as Customer Type, Family size, and location identifiers.
  • Predictive Modeling — compared Logistic Regression and Random Forest using 5-Fold Stratified Cross-Validation.
  • Threshold Optimization — tuned the churn decision threshold using Out-of-Fold predictions with a minimum 40% Churn Precision constraint.
  • Deployment — integrated the final model into a Streamlit application for individual and batch customer prediction.

Conclusion.

The project combined customer analytics and predictive modeling to turn churn data into actionable retention insights. The analysis identified meaningful differences across feedback, customer frequency, demographics, occupation, and income, while the machine learning stage produced a Random Forest model with 76.5% Churn Recall, 61.9% Churn Precision, and 0.834 ROC-AUC on the holdout test set. The final solution was deployed through a Streamlit application, enabling individual and batch customer risk predictions. The project demonstrates an end-to-end workflow from exploratory analysis to predictive decision support for customer retention.