Mathew K Analytics

Lesson 45 · Data Mining

Comprehensive Data Mining Capstone Project: Step-by-Step Case Study Walkthrough

Welcome to your final project! In this lesson, you will combine everything you have learned to analyze and model real datasets. We will work with data about…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Week 12 Capstone Project - Comprehensive Data Mining Case Study#

Welcome to your final project! In this lesson, you will combine everything you have learned to analyze and model real datasets. We will work with data about telecom customer churn, credit card fraud, and mall customers.

Let us dive in and put your data mining skills to the test!

# Suppress all warnings for a clean output
import warnings; warnings.filterwarnings("ignore")
import numpy as np
np.random.seed(42)

Part 1: Data Setup#

For our capstone, we will use the Telecom Customer Churn dataset. This data includes information about customers, their services, and whether they left (churned). Churn analysis is important for businesses to understand why customers leave.

# Data setup (Telecom Customer Churn Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/IBM/telco-customer-churn-on-icp4d/master/data/Telco-Customer-Churn.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
(7043, 21)
   customerID  gender  SeniorCitizen Partner Dependents  tenure PhoneService  \
0  7590-VHVEG  Female              0     Yes         No       1           No   
1  5575-GNVDE    Male              0      No         No      34          Yes   
2  3668-QPYBK    Male              0      No         No       2          Yes   

      MultipleLines InternetService OnlineSecurity  ... DeviceProtection  \
0  No phone service             DSL             No  ...               No   
1                No             DSL            Yes  ...              Yes   
2                No             DSL            Yes  ...               No   

  TechSupport StreamingTV StreamingMovies        Contract PaperlessBilling  \
0          No          No              No  Month-to-month              Yes   
1          No          No              No        One year               No   
2          No          No              No  Month-to-month              Yes   

      PaymentMethod MonthlyCharges  TotalCharges Churn  
0  Electronic check          29.85         29.85    No  
1      Mailed check          56.95        1889.5    No  
2      Mailed check          53.85        108.15   Yes  

[3 rows x 21 columns]
# Quick look at missing data
print(df.isnull().sum().sort_values(ascending=False).head(8))
customerID       0
gender           0
SeniorCitizen    0
Partner          0
Dependents       0
tenure           0
PhoneService     0
MultipleLines    0
dtype: int64
# Handle missing data simply
df = df.dropna()
print(df.shape)
(7043, 21)

Part 2: Data Preprocessing#

Let us convert text columns to numbers. Machine learning models in Python need free of text values. We will change categories into numerical codes.

# Convert text columns to codes
for col in df.select_dtypes('object').columns:
    df[col] = df[col].astype('category').cat.codes
print(df.head(3))
   customerID  gender  SeniorCitizen  Partner  Dependents  tenure  \
0        5375       0              0        1           0       1   
1        3962       1              0        0           0      34   
2        2564       1              0        0           0       2   

   PhoneService  MultipleLines  InternetService  OnlineSecurity  ...  \
0             0              1                0               0  ...   
1             1              0                0               2  ...   
2             1              0                0               2  ...   

   DeviceProtection  TechSupport  StreamingTV  StreamingMovies  Contract  \
0                 0            0            0                0         0   
1                 2            0            0                0         1   
2                 0            0            0                0         0   

   PaperlessBilling  PaymentMethod  MonthlyCharges  TotalCharges  Churn  
0                 1              2           29.85          2505      0  
1                 0              3           56.95          1466      0  
2                 1              3           53.85           157      1  

[3 rows x 21 columns]

Part 3: Exploratory Data Analysis (EDA)#

Let us get to know our data by looking at descriptive statistics and class balance. Exploring data helps us spot problems and interesting patterns.

# Show summary statistics
print(df.describe().T)
                   count         mean          std    min     25%      50%  \
customerID        7043.0  3521.000000  2033.283305   0.00  1760.5  3521.00   
gender            7043.0     0.504756     0.500013   0.00     0.0     1.00   
SeniorCitizen     7043.0     0.162147     0.368612   0.00     0.0     0.00   
Partner           7043.0     0.483033     0.499748   0.00     0.0     0.00   
Dependents        7043.0     0.299588     0.458110   0.00     0.0     0.00   
tenure            7043.0    32.371149    24.559481   0.00     9.0    29.00   
PhoneService      7043.0     0.903166     0.295752   0.00     1.0     1.00   
MultipleLines     7043.0     0.940508     0.948554   0.00     0.0     1.00   
InternetService   7043.0     0.872923     0.737796   0.00     0.0     1.00   
OnlineSecurity    7043.0     0.790004     0.859848   0.00     0.0     1.00   
OnlineBackup      7043.0     0.906432     0.880162   0.00     0.0     1.00   
DeviceProtection  7043.0     0.904444     0.879949   0.00     0.0     1.00   
TechSupport       7043.0     0.797104     0.861551   0.00     0.0     1.00   
StreamingTV       7043.0     0.985376     0.885002   0.00     0.0     1.00   
StreamingMovies   7043.0     0.992475     0.885091   0.00     0.0     1.00   
Contract          7043.0     0.690473     0.833755   0.00     0.0     0.00   
PaperlessBilling  7043.0     0.592219     0.491457   0.00     0.0     1.00   
PaymentMethod     7043.0     1.574329     1.068104   0.00     1.0     2.00   
MonthlyCharges    7043.0    64.761692    30.090047  18.25    35.5    70.35   
TotalCharges      7043.0  3257.794122  1888.693496   0.00  1609.0  3249.00   
Churn             7043.0     0.265370     0.441561   0.00     0.0     0.00   

                      75%      max  
customerID        5281.50  7042.00  
gender               1.00     1.00  
SeniorCitizen        0.00     1.00  
Partner              1.00     1.00  
Dependents           1.00     1.00  
tenure              55.00    72.00  
PhoneService         1.00     1.00  
MultipleLines        2.00     2.00  
InternetService      1.00     2.00  
OnlineSecurity       2.00     2.00  
OnlineBackup         2.00     2.00  
DeviceProtection     2.00     2.00  
TechSupport          2.00     2.00  
StreamingTV          2.00     2.00  
StreamingMovies      2.00     2.00  
Contract             1.00     2.00  
PaperlessBilling     1.00     1.00  
PaymentMethod        2.00     3.00  
MonthlyCharges      89.85   118.75  
TotalCharges      4901.50  6530.00  
Churn                1.00     1.00  
# Class balance for churn label
print(df['Churn'].value_counts())
Churn
0    5174
1    1869
Name: count, dtype: int64

Part 4: Association Rule Mining (Mini-Project)#

Now we will use market basket analysis to discover products that are often bought together. Let us use the Groceries Dataset for this step.

# Data setup (Groceries Dataset)
import pandas as pd
import kagglehub, os
path = kagglehub.dataset_download('heeraldedhia/groceries-dataset')
files = os.listdir(path)
csv_file = [f for f in files if f.endswith('.csv')][0]
df_groceries = pd.read_csv(os.path.join(path, csv_file))
print(df_groceries.shape)
print(df_groceries.head(3))
(38765, 3)
   Member_number        Date itemDescription
0           1808  21-07-2015  tropical fruit
1           2552  05-01-2015      whole milk
2           2300  19-09-2015       pip fruit
# Prepare basket data for association rules
transactions = df_groceries.groupby('Member_number')['itemDescription'].apply(list).values.tolist()
print(transactions[:2])
[['soda', 'canned beer', 'sausage', 'sausage', 'whole milk', 'whole milk', 'pickled vegetables', 'misc. beverages', 'semi-finished bread', 'hygiene articles', 'yogurt', 'pastry', 'salty snack'], ['frankfurter', 'frankfurter', 'beef', 'sausage', 'whole milk', 'soda', 'curd', 'white bread', 'whole milk', 'soda', 'whipped/sour cream', 'rolls/buns']]
# Simple association rules with mlxtend
from mlxtend.preprocessing import TransactionEncoder
from mlxtend.frequent_patterns import apriori, association_rules

# Encode baskets for mining
te = TransactionEncoder()
te_ary = te.fit(transactions).transform(transactions)
basket_df = pd.DataFrame(te_ary, columns=te.columns_)

# Find frequent itemsets
frequent = apriori(basket_df, min_support=0.04, use_colnames=True)
rules = association_rules(frequent, metric='lift', min_threshold=1.2)
print(rules[['antecedents', 'consequents', 'support', 'confidence', 'lift']].head())
          antecedents         consequents   support  confidence      lift
0  (other vegetables)            (butter)  0.057209    0.151907  1.201085
1            (butter)  (other vegetables)  0.057209    0.452333  1.201085
2            (yogurt)            (butter)  0.044895    0.158658  1.254462
3            (butter)            (yogurt)  0.044895    0.354970  1.254462
4   (root vegetables)       (canned beer)  0.047460    0.205784  1.245570

Part 5: Classification Modeling#

It's time to use a machine learning model to predict if a customer will churn. We will use the Random Forest classifier for this task.

# Split data into features and labels and create train/test sets
from sklearn.model_selection import train_test_split
X = df.drop('Churn', axis=1)
y = df['Churn']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
print(X_train.shape, X_test.shape)
(5634, 20) (1409, 20)
# Train a Random Forest Classifier
from sklearn.ensemble import RandomForestClassifier
rf = RandomForestClassifier(n_estimators=100, random_state=42)
rf.fit(X_train, y_train)
print('Training done!')
Training done!
# Check accuracy on the test set
score = rf.score(X_test, y_test)
print(f'Accuracy: {score:.3f}')
Accuracy: 0.797

Part 6: Clustering and Visualization#

For our next analysis, let us try clustering on the Mall Customers Dataset. Clustering finds natural groups in data, even without labels.

# Data setup (Mall Customers Dataset)
import pandas as pd
url = 'https://gist.githubusercontent.com/pravalliyaram/5c05f43d2351249927b8a3f3cc3e5ecf/raw/Mall_Customers.csv'
mall = pd.read_csv(url)
print(mall.shape)
print(mall.head(3))
(200, 5)
   CustomerID  Gender  Age  Annual Income (k$)  Spending Score (1-100)
0           1    Male   19                  15                      39
1           2    Male   21                  15                      81
2           3  Female   20                  16                       6
# KMeans clustering on 'Annual Income' and 'Spending Score'
from sklearn.cluster import KMeans
import matplotlib.pyplot as plt

X = mall[['Annual Income (k$)', 'Spending Score (1-100)']]
kmeans = KMeans(n_clusters=4, random_state=42)
clusters = kmeans.fit_predict(X)
mall['Cluster'] = clusters
plt.scatter(X['Annual Income (k$)'], X['Spending Score (1-100)'], c=clusters, cmap='tab10')
plt.xlabel('Annual Income (k$)')
plt.ylabel('Spending Score (1-100)')
plt.title('Mall Customers Clusters')
plt.show()
No description has been provided for this image

Part 7: Advanced Topics (Ensemble & Anomaly Detection)#

Ensembles combine several models to boost accuracy. Anomaly detection helps find fraud or rare events.

# Anomaly detection: Credit Card Fraud Dataset
import pandas as pd
url = 'https://storage.googleapis.com/download.tensorflow.org/data/creditcard.csv'
df_fraud = pd.read_csv(url)
print(df_fraud.shape)
print(df_fraud['Class'].value_counts())
(284807, 31)
Class
0    284315
1       492
Name: count, dtype: int64
# Quick anomaly detection with Isolation Forest
from sklearn.ensemble import IsolationForest
features = df_fraud.drop(['Class'], axis=1).sample(10000, random_state=42)
isf = IsolationForest(contamination=0.005, random_state=42)
anomaly_pred = isf.fit_predict(features)
print('Number of predicted anomalies:', sum(anomaly_pred == -1))
Number of predicted anomalies: 50

Part 8: Time Series Basics#

Time series data helps us spot trends over time, useful for business and forecasting. Here is a tiny example with made-up data just to show the idea.

# Tiny time series example
import pandas as pd
import matplotlib.pyplot as plt
ts = pd.Series([230, 245, 260, 265, 275, 290, 300],
              index=pd.date_range('2023-01-31', periods=7, freq='M'))
ts.plot(marker='o')
plt.title('Sample Customer Count Over Months')
plt.ylabel('Customer Count')
plt.show()
No description has been provided for this image

Part 9: Mini-Project Create Your Own Churn Predictor!#

Let us put it all together. Build a model to predict which customers might leave. Try it on a small sample from the telecom data.

# Enter a sample customer for prediction
print('Enter customer features separated by commas (see column info):')
custom_input = input()
features = list(map(int, custom_input.strip().split(',')))
pred = rf.predict([features])[0]
print('Predicted churn:' if pred == 1 else 'Predicted stay!')
Enter customer features separated by commas (see column info):
Predicted stay!
# Challenge: Try changing the churn model to use logistic regression
from sklearn.linear_model import LogisticRegression
logreg = LogisticRegression(max_iter=200, random_state=42)
logreg.fit(X_train, y_train)
lr_score = logreg.score(X_test, y_test)
print(f'Logistic Regression Accuracy: {lr_score:.3f}')
Logistic Regression Accuracy: 0.813

Part 10: Best Practices & Final Tips#

When working with real data, always:

  • Check your data for quality and missing values
  • Split into train and test sets
  • Try multiple models and compare
  • Tune hyperparameters for even better results

Lastly: Practice, experiment, and keep learning!

Congratulations! You completed the Data Mining Capstone.#

Share your progress and questions in the comments. For more tutorials, subscribe and hit the notification bell! Good luck and keep exploring data!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.