Lesson 45 · Data Mining
Comprehensive Data Mining Capstone Project: Step-by-Step Case Study Walkthrough
Welcome to your final project! In this lesson, you will combine everything you have learned to analyze and model real datasets. We will work with data about…
- CourseData Mining
- Lesson45 of 31
- Video26 min
- FormatJupyter notebook · 20 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbWeek 12 Capstone Project - Comprehensive Data Mining Case Study#
Welcome to your final project! In this lesson, you will combine everything you have learned to analyze and model real datasets. We will work with data about telecom customer churn, credit card fraud, and mall customers.
Let us dive in and put your data mining skills to the test!
# Suppress all warnings for a clean output
import warnings; warnings.filterwarnings("ignore")
import numpy as np
np.random.seed(42)
Part 1: Data Setup#
For our capstone, we will use the Telecom Customer Churn dataset. This data includes information about customers, their services, and whether they left (churned). Churn analysis is important for businesses to understand why customers leave.
# Data setup (Telecom Customer Churn Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/IBM/telco-customer-churn-on-icp4d/master/data/Telco-Customer-Churn.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
# Quick look at missing data
print(df.isnull().sum().sort_values(ascending=False).head(8))
# Handle missing data simply
df = df.dropna()
print(df.shape)
Part 2: Data Preprocessing#
Let us convert text columns to numbers. Machine learning models in Python need free of text values. We will change categories into numerical codes.
# Convert text columns to codes
for col in df.select_dtypes('object').columns:
df[col] = df[col].astype('category').cat.codes
print(df.head(3))
Part 3: Exploratory Data Analysis (EDA)#
Let us get to know our data by looking at descriptive statistics and class balance. Exploring data helps us spot problems and interesting patterns.
# Show summary statistics
print(df.describe().T)
# Class balance for churn label
print(df['Churn'].value_counts())
Part 4: Association Rule Mining (Mini-Project)#
Now we will use market basket analysis to discover products that are often bought together. Let us use the Groceries Dataset for this step.
# Data setup (Groceries Dataset)
import pandas as pd
import kagglehub, os
path = kagglehub.dataset_download('heeraldedhia/groceries-dataset')
files = os.listdir(path)
csv_file = [f for f in files if f.endswith('.csv')][0]
df_groceries = pd.read_csv(os.path.join(path, csv_file))
print(df_groceries.shape)
print(df_groceries.head(3))
# Prepare basket data for association rules
transactions = df_groceries.groupby('Member_number')['itemDescription'].apply(list).values.tolist()
print(transactions[:2])
# Simple association rules with mlxtend
from mlxtend.preprocessing import TransactionEncoder
from mlxtend.frequent_patterns import apriori, association_rules
# Encode baskets for mining
te = TransactionEncoder()
te_ary = te.fit(transactions).transform(transactions)
basket_df = pd.DataFrame(te_ary, columns=te.columns_)
# Find frequent itemsets
frequent = apriori(basket_df, min_support=0.04, use_colnames=True)
rules = association_rules(frequent, metric='lift', min_threshold=1.2)
print(rules[['antecedents', 'consequents', 'support', 'confidence', 'lift']].head())
Part 5: Classification Modeling#
It's time to use a machine learning model to predict if a customer will churn. We will use the Random Forest classifier for this task.
# Split data into features and labels and create train/test sets
from sklearn.model_selection import train_test_split
X = df.drop('Churn', axis=1)
y = df['Churn']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
print(X_train.shape, X_test.shape)
# Train a Random Forest Classifier
from sklearn.ensemble import RandomForestClassifier
rf = RandomForestClassifier(n_estimators=100, random_state=42)
rf.fit(X_train, y_train)
print('Training done!')
# Check accuracy on the test set
score = rf.score(X_test, y_test)
print(f'Accuracy: {score:.3f}')
Part 6: Clustering and Visualization#
For our next analysis, let us try clustering on the Mall Customers Dataset. Clustering finds natural groups in data, even without labels.
# Data setup (Mall Customers Dataset)
import pandas as pd
url = 'https://gist.githubusercontent.com/pravalliyaram/5c05f43d2351249927b8a3f3cc3e5ecf/raw/Mall_Customers.csv'
mall = pd.read_csv(url)
print(mall.shape)
print(mall.head(3))
# KMeans clustering on 'Annual Income' and 'Spending Score'
from sklearn.cluster import KMeans
import matplotlib.pyplot as plt
X = mall[['Annual Income (k$)', 'Spending Score (1-100)']]
kmeans = KMeans(n_clusters=4, random_state=42)
clusters = kmeans.fit_predict(X)
mall['Cluster'] = clusters
plt.scatter(X['Annual Income (k$)'], X['Spending Score (1-100)'], c=clusters, cmap='tab10')
plt.xlabel('Annual Income (k$)')
plt.ylabel('Spending Score (1-100)')
plt.title('Mall Customers Clusters')
plt.show()
Part 7: Advanced Topics (Ensemble & Anomaly Detection)#
Ensembles combine several models to boost accuracy. Anomaly detection helps find fraud or rare events.
# Anomaly detection: Credit Card Fraud Dataset
import pandas as pd
url = 'https://storage.googleapis.com/download.tensorflow.org/data/creditcard.csv'
df_fraud = pd.read_csv(url)
print(df_fraud.shape)
print(df_fraud['Class'].value_counts())
# Quick anomaly detection with Isolation Forest
from sklearn.ensemble import IsolationForest
features = df_fraud.drop(['Class'], axis=1).sample(10000, random_state=42)
isf = IsolationForest(contamination=0.005, random_state=42)
anomaly_pred = isf.fit_predict(features)
print('Number of predicted anomalies:', sum(anomaly_pred == -1))
Part 8: Time Series Basics#
Time series data helps us spot trends over time, useful for business and forecasting. Here is a tiny example with made-up data just to show the idea.
# Tiny time series example
import pandas as pd
import matplotlib.pyplot as plt
ts = pd.Series([230, 245, 260, 265, 275, 290, 300],
index=pd.date_range('2023-01-31', periods=7, freq='M'))
ts.plot(marker='o')
plt.title('Sample Customer Count Over Months')
plt.ylabel('Customer Count')
plt.show()
Part 9: Mini-Project Create Your Own Churn Predictor!#
Let us put it all together. Build a model to predict which customers might leave. Try it on a small sample from the telecom data.
# Enter a sample customer for prediction
print('Enter customer features separated by commas (see column info):')
custom_input = input()
features = list(map(int, custom_input.strip().split(',')))
pred = rf.predict([features])[0]
print('Predicted churn:' if pred == 1 else 'Predicted stay!')
# Challenge: Try changing the churn model to use logistic regression
from sklearn.linear_model import LogisticRegression
logreg = LogisticRegression(max_iter=200, random_state=42)
logreg.fit(X_train, y_train)
lr_score = logreg.score(X_test, y_test)
print(f'Logistic Regression Accuracy: {lr_score:.3f}')
Part 10: Best Practices & Final Tips#
When working with real data, always:
- Check your data for quality and missing values
- Split into train and test sets
- Try multiple models and compare
- Tune hyperparameters for even better results
Lastly: Practice, experiment, and keep learning!
Congratulations! You completed the Data Mining Capstone.#
Share your progress and questions in the comments. For more tutorials, subscribe and hit the notification bell! Good luck and keep exploring data!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



