Mathew K Analytics

Lesson 30 · Data analytics zero to hero

End-to-End Analytics Capstone Project | Data Analytics #30

The final video of the 30-part series. One complete, real project, start to finish, drawing on everything from variables in video one to machine learning in…

What you'll learn

Datasets used in this lesson

Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.

📓 Full notebook

Download .ipynb

Data Analytics Zero to Hero, Video 30: Capstone, End-to-End Data Analyst Project#

  • The final video of the 30-part series. One complete, real project, start to finish, drawing on everything from variables in video one to machine learning in video twenty-six.
  • A brand new real dataset for this final video: the classic Mall Customer Segmentation dataset, 200 real shoppers with real age, income, and spending data.
  • The business question: what are the genuine customer segments in this data, and how should each be treated differently?
  • Let's jump straight in.

Before You Start#

  • Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
  • Place mall_customers_clustered.csv in the same folder as this notebook.
  • Everything you need, pandas, NumPy, Matplotlib, and scikit-learn, has already been installed across this series.

Part 1: Load and Clean#

import pandas as pd
import numpy as np

df = pd.read_csv('mall_customers_clustered.csv')
df = df.drop(columns='Cluster')
print(df.shape)
print(df.isna().sum().sum(), 'real missing values')
print(df.head())
(200, 5)
0 real missing values
   CustomerID  Gender  Age  Annual Income (k$)  Spending Score (1-100)
0           1    Male   19                  15                      39
1           2    Male   21                  15                      81
2           3  Female   20                  16                       6
3           4  Female   23                  16                      77
4           5  Female   31                  17                      40

Part 2: Exploratory Analysis#

print(df.describe())
print(df['Gender'].value_counts())
       CustomerID         Age  Annual Income (k$)  Spending Score (1-100)
count  200.000000  200.000000          200.000000              200.000000
mean   100.500000   38.850000           60.560000               50.200000
std     57.879185   13.969007           26.264721               25.823522
min      1.000000   18.000000           15.000000                1.000000
25%     50.750000   28.750000           41.500000               34.750000
50%    100.500000   36.000000           61.500000               50.000000
75%    150.250000   49.000000           78.000000               73.000000
max    200.000000   70.000000          137.000000               99.000000
Gender
Female    112
Male       88
Name: count, dtype: int64
numeric_cols = ['Age', 'Annual Income (k$)', 'Spending Score (1-100)']
print(df[numeric_cols].corr().round(2))
                         Age  Annual Income (k$)  Spending Score (1-100)
Age                     1.00               -0.01                   -0.33
Annual Income (k$)     -0.01                1.00                    0.01
Spending Score (1-100) -0.33                0.01                    1.00

Part 3: Visualizing the Real Relationship#

import matplotlib.pyplot as plt

plt.figure(figsize=(7, 5))
plt.scatter(df['Annual Income (k$)'], df['Spending Score (1-100)'], alpha=0.7)
plt.title('Real Annual Income vs. Spending Score')
plt.xlabel('Annual Income (k$)')
plt.ylabel('Spending Score (1-100)')
plt.show()
No description has been provided for this image

Part 4: Does Gender Affect Spending?#

from scipy import stats

male_scores = df[df['Gender'] == 'Male']['Spending Score (1-100)']
female_scores = df[df['Gender'] == 'Female']['Spending Score (1-100)']
t_stat, p_val = stats.ttest_ind(male_scores, female_scores)
print(f'Male mean: {male_scores.mean():.1f}, Female mean: {female_scores.mean():.1f}')
print(f't-statistic: {t_stat:.3f}, p-value: {p_val:.3f}')
Male mean: 48.5, Female mean: 51.5
t-statistic: -0.819, p-value: 0.414

Part 5: Finding Real Segments with K-Means#

from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans

X = df[['Annual Income (k$)', 'Spending Score (1-100)']]
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)

inertias = []
for k in range(1, 9):
    km = KMeans(n_clusters=k, random_state=42, n_init=10)
    km.fit(X_scaled)
    inertias.append(km.inertia_)
print([round(i, 1) for i in inertias])
[400.0, 269.7, 157.7, 108.9, 65.6, 55.1, 44.9, 37.2]
final_km = KMeans(n_clusters=5, random_state=42, n_init=10)
df['segment'] = final_km.fit_predict(X_scaled)
print(df['segment'].value_counts().sort_index())
print(df.groupby('segment')[['Annual Income (k$)', 'Spending Score (1-100)']].mean().round(1))
segment
0    81
1    39
2    22
3    35
4    23
Name: count, dtype: int64
         Annual Income (k$)  Spending Score (1-100)
segment                                            
0                      55.3                    49.5
1                      86.5                    82.1
2                      25.7                    79.4
3                      88.2                    17.1
4                      26.3                    20.9
plt.figure(figsize=(7, 5))
plt.scatter(df['Annual Income (k$)'], df['Spending Score (1-100)'], c=df['segment'], cmap='tab10')
plt.title('Real Customer Segments from K-Means')
plt.xlabel('Annual Income (k$)')
plt.ylabel('Spending Score (1-100)')
plt.show()
No description has been provided for this image

Part 6: Turning Clusters into a Real Business Story#

def label_segment(row):
    income, spend = row['Annual Income (k$)'], row['Spending Score (1-100)']
    if income > 60 and spend > 60:
        return 'High Income, High Spend'
    elif income > 60 and spend <= 60:
        return 'High Income, Low Spend'
    elif income <= 60 and spend > 60:
        return 'Low Income, High Spend'
    else:
        return 'Low Income, Low Spend'

segment_means = df.groupby('segment')[['Annual Income (k$)', 'Spending Score (1-100)']].mean()
segment_names = segment_means.apply(label_segment, axis=1)
print(segment_names)
segment
0      Low Income, Low Spend
1    High Income, High Spend
2     Low Income, High Spend
3     High Income, Low Spend
4      Low Income, Low Spend
dtype: object

Part 7: A Final Real Report#

biggest_segment = df['segment'].value_counts().idxmax()
biggest_size = df['segment'].value_counts().max()
biggest_label = segment_names[biggest_segment]

summary = f'''
CUSTOMER SEGMENTATION SUMMARY
K-means identified 5 genuine customer segments among {len(df)} real shoppers.
The largest segment, '{biggest_label}', contains {biggest_size} customers ({biggest_size / len(df):.0%} of the total).
Gender alone does not significantly predict spending score (p={p_val:.2f}), so segmentation should rely on income and spending behavior instead.
Recommendation: target 'High Income, Low Spend' customers with personalized offers, and reward 'High Income, High Spend' customers with loyalty perks to protect that real revenue.
'''
print(summary)
CUSTOMER SEGMENTATION SUMMARY
K-means identified 5 genuine customer segments among 200 real shoppers.
The largest segment, 'Low Income, Low Spend', contains 81 customers (40% of the total).
Gender alone does not significantly predict spending score (p=0.41), so segmentation should rely on income and spending behavior instead.
Recommendation: target 'High Income, Low Spend' customers with personalized offers, and reward 'High Income, High Spend' customers with loyalty perks to protect that real revenue.

Series Wrap-Up: Data Analytics Zero to Hero#

  • Video one to five: real Python foundations, variables through file handling.
  • Video six to twelve: NumPy and a deep, multi-part pandas block, cleaning, wrangling, grouping, merging, and time series, all on real data.
  • Video thirteen to fifteen: visualization with Matplotlib and Seaborn, and a real interactive Streamlit dashboard.
  • Video sixteen to eighteen: SQL fundamentals, joins, and advanced analyst queries, entirely through Python's sqlite3.
  • Video nineteen and twenty: real Excel automation with openpyxl, formulas, formatting, pivots, and native charts.
  • Video twenty-one and twenty-two: connecting to live APIs and scraping real web data.
  • Video twenty-three and twenty-four: statistics, from descriptive measures to hypothesis testing.
  • Video twenty-five and twenty-six: your first real machine learning models, evaluated properly.
  • Video twenty-seven: turning analysis into a real, effective story.
  • Video twenty-eight to thirty: three complete, real end-to-end projects, tying every single skill together.
  • That's the whole series: zero to hero, entirely in Python, entirely on real data. Thank you for building all thirty of these with me. If this helped, subscribing and leaving a comment genuinely helps this channel keep making more.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.