Lesson 22 · Data Mining
Customer Churn Prediction: Step-by-Step Python Tutorial Using Real Data
Welcome! This tutorial will guide you through a real-world data mining project: predicting customer churn with Python. You will learn step by step from data…
- CourseData Mining
- Lesson22 of 31
- Video23 min
- FormatJupyter notebook · 14 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbWeek 56 Hands-on Project: Predicting Customer Churn#
Welcome! This tutorial will guide you through a real-world data mining project: predicting customer churn with Python.
You will learn step by step from data loading to model evaluation.
Let us get started!
import warnings
import numpy as np
np.random.seed(42)
warnings.filterwarnings("ignore")
# Suppress any warnings for a smoother experience
What is Customer Churn?#
Churn means when customers leave a service or company. Businesses want to predict churn so they can prevent it.
Data mining helps us find patterns and predict who may leave.
# Data setup (Telecom Customer Churn Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/IBM/telco-customer-churn-on-icp4d/master/data/Telco-Customer-Churn.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
Step 1: Data Exploration#
First, let us see what features the dataset has.
Understanding your data is always the first step.
print(df.columns)
# Print the column names
df.info()
# Show summary info, including nulls
Step 2: Data Cleaning#
Good models need clean data.
Let us handle missing values and tidy up columns for easier use.
# Quick cleaning: drop 'customerID', fill missing 'TotalCharges'
df = df.drop('customerID', axis=1)
df['TotalCharges'] = pd.to_numeric(df['TotalCharges'], errors='coerce')
df['TotalCharges'].fillna(df['TotalCharges'].median(), inplace=True)
# Encode categorical variables for modeling
df_encoded = pd.get_dummies(df, drop_first=True)
print(df_encoded.shape)
Step 3: Quick Data Visualization#
Let us look at the distribution of churned versus non-churned customers.
Visualization helps us spot imbalances and trends.
import matplotlib.pyplot as plt
df['Churn'].value_counts().plot(kind='bar', color=['skyblue', 'salmon'])
plt.title('Churn vs Non-Churn Customers')
plt.xlabel('Churn')
plt.ylabel('Count')
plt.show()
Step 4: Splitting Data for Training and Testing#
We must train on some data and test on new data to check our model.
Let us split the data now.
from sklearn.model_selection import train_test_split
X = df_encoded.drop('Churn_Yes', axis=1)
y = df_encoded['Churn_Yes']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
Step 5: Training Our First Classification Model#
Let us use a decision tree for a first prediction.
This will help us see basic accuracy and set a baseline.
from sklearn.tree import DecisionTreeClassifier
clf = DecisionTreeClassifier(random_state=42)
clf.fit(X_train, y_train)
y_pred = clf.predict(X_test)
from sklearn.metrics import accuracy_score
accuracy = accuracy_score(y_test, y_pred)
print(f"Accuracy: {accuracy:.2f}")
# Try logistic regression for comparison
from sklearn.linear_model import LogisticRegression
lr = LogisticRegression(max_iter=1000, random_state=42)
lr.fit(X_train, y_train)
y_pred_lr = lr.predict(X_test)
print(f"Logistic Regression Accuracy: {accuracy_score(y_test, y_pred_lr):.2f}")
# Confusion matrix to show more detail
from sklearn.metrics import confusion_matrix
import seaborn as sns
cm = confusion_matrix(y_test, y_pred)
sns.heatmap(cm, annot=True, fmt='d', cmap='Blues')
plt.xlabel('Predicted')
plt.ylabel('Actual')
plt.title('Decision Tree Confusion Matrix')
plt.show()
Step 6: Feature Importance#
Knowing which features are most important helps businesses act.
Let us look at the top drivers of churn from our tree.
import numpy as np
# Get top 5 features
feat_import = pd.Series(clf.feature_importances_, index=X_train.columns)
feat_import.nlargest(5).plot(kind='barh', color='purple')
plt.title('Top 5 Important Features for Churn')
plt.show()
Step 7: Try a Mini Challenge#
Now it is your turn! Change the model or features and see how accuracy changes.
Extra: Try adding a new column or dropping one.
# Practice: Try input() to choose test size
test_size = float(input("Type a test size fraction (like 0.3 for 30% test): "))
X_train2, X_test2, y_train2, y_test2 = train_test_split(X, y, test_size=test_size, random_state=42)
clf2 = DecisionTreeClassifier(random_state=42)
clf2.fit(X_train2, y_train2)
print("New accuracy:", accuracy_score(y_test2, clf2.predict(X_test2)))
Recap: What We Learned#
- Load and explore churn data
- Clean and encode features
- Train classifiers and check accuracy
- Visualize feature importance
Learning by doing is the key. Well done!
Next Steps and Extra Tips#
- Try other classifiers like Random Forest or XGBoost
- Explore cross-validation for better reliability
- Always visualize and check your predictions
Remember, practice is powerful!
Thank you for learning with us!
If you found this helpful, please subscribe and share the video.
See you in the next lesson!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



