Lesson 30 · Data analytics zero to hero
End-to-End Analytics Capstone Project | Data Analytics #30
The final video of the 30-part series. One complete, real project, start to finish, drawing on everything from variables in video one to machine learning in…
- CourseData analytics zero to hero
- Lesson30 of 30
- Video15 min
- FormatJupyter notebook · 10 code cells
- Data1 dataset
What you'll learn
Datasets used in this lesson
Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.
📓 Full notebook
Download .ipynbData Analytics Zero to Hero, Video 30: Capstone, End-to-End Data Analyst Project#
- The final video of the 30-part series. One complete, real project, start to finish, drawing on everything from variables in video one to machine learning in video twenty-six.
- A brand new real dataset for this final video: the classic Mall Customer Segmentation dataset, 200 real shoppers with real age, income, and spending data.
- The business question: what are the genuine customer segments in this data, and how should each be treated differently?
- Let's jump straight in.
Before You Start#
- Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
- Place mall_customers_clustered.csv in the same folder as this notebook.
- Everything you need, pandas, NumPy, Matplotlib, and scikit-learn, has already been installed across this series.
Part 1: Load and Clean#
import pandas as pd
import numpy as np
df = pd.read_csv('mall_customers_clustered.csv')
df = df.drop(columns='Cluster')
print(df.shape)
print(df.isna().sum().sum(), 'real missing values')
print(df.head())
Part 2: Exploratory Analysis#
print(df.describe())
print(df['Gender'].value_counts())
numeric_cols = ['Age', 'Annual Income (k$)', 'Spending Score (1-100)']
print(df[numeric_cols].corr().round(2))
Part 3: Visualizing the Real Relationship#
import matplotlib.pyplot as plt
plt.figure(figsize=(7, 5))
plt.scatter(df['Annual Income (k$)'], df['Spending Score (1-100)'], alpha=0.7)
plt.title('Real Annual Income vs. Spending Score')
plt.xlabel('Annual Income (k$)')
plt.ylabel('Spending Score (1-100)')
plt.show()
Part 4: Does Gender Affect Spending?#
from scipy import stats
male_scores = df[df['Gender'] == 'Male']['Spending Score (1-100)']
female_scores = df[df['Gender'] == 'Female']['Spending Score (1-100)']
t_stat, p_val = stats.ttest_ind(male_scores, female_scores)
print(f'Male mean: {male_scores.mean():.1f}, Female mean: {female_scores.mean():.1f}')
print(f't-statistic: {t_stat:.3f}, p-value: {p_val:.3f}')
Part 5: Finding Real Segments with K-Means#
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
X = df[['Annual Income (k$)', 'Spending Score (1-100)']]
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)
inertias = []
for k in range(1, 9):
km = KMeans(n_clusters=k, random_state=42, n_init=10)
km.fit(X_scaled)
inertias.append(km.inertia_)
print([round(i, 1) for i in inertias])
final_km = KMeans(n_clusters=5, random_state=42, n_init=10)
df['segment'] = final_km.fit_predict(X_scaled)
print(df['segment'].value_counts().sort_index())
print(df.groupby('segment')[['Annual Income (k$)', 'Spending Score (1-100)']].mean().round(1))
plt.figure(figsize=(7, 5))
plt.scatter(df['Annual Income (k$)'], df['Spending Score (1-100)'], c=df['segment'], cmap='tab10')
plt.title('Real Customer Segments from K-Means')
plt.xlabel('Annual Income (k$)')
plt.ylabel('Spending Score (1-100)')
plt.show()
Part 6: Turning Clusters into a Real Business Story#
def label_segment(row):
income, spend = row['Annual Income (k$)'], row['Spending Score (1-100)']
if income > 60 and spend > 60:
return 'High Income, High Spend'
elif income > 60 and spend <= 60:
return 'High Income, Low Spend'
elif income <= 60 and spend > 60:
return 'Low Income, High Spend'
else:
return 'Low Income, Low Spend'
segment_means = df.groupby('segment')[['Annual Income (k$)', 'Spending Score (1-100)']].mean()
segment_names = segment_means.apply(label_segment, axis=1)
print(segment_names)
Part 7: A Final Real Report#
biggest_segment = df['segment'].value_counts().idxmax()
biggest_size = df['segment'].value_counts().max()
biggest_label = segment_names[biggest_segment]
summary = f'''
CUSTOMER SEGMENTATION SUMMARY
K-means identified 5 genuine customer segments among {len(df)} real shoppers.
The largest segment, '{biggest_label}', contains {biggest_size} customers ({biggest_size / len(df):.0%} of the total).
Gender alone does not significantly predict spending score (p={p_val:.2f}), so segmentation should rely on income and spending behavior instead.
Recommendation: target 'High Income, Low Spend' customers with personalized offers, and reward 'High Income, High Spend' customers with loyalty perks to protect that real revenue.
'''
print(summary)
Series Wrap-Up: Data Analytics Zero to Hero#
- Video one to five: real Python foundations, variables through file handling.
- Video six to twelve: NumPy and a deep, multi-part pandas block, cleaning, wrangling, grouping, merging, and time series, all on real data.
- Video thirteen to fifteen: visualization with Matplotlib and Seaborn, and a real interactive Streamlit dashboard.
- Video sixteen to eighteen: SQL fundamentals, joins, and advanced analyst queries, entirely through Python's sqlite3.
- Video nineteen and twenty: real Excel automation with openpyxl, formulas, formatting, pivots, and native charts.
- Video twenty-one and twenty-two: connecting to live APIs and scraping real web data.
- Video twenty-three and twenty-four: statistics, from descriptive measures to hypothesis testing.
- Video twenty-five and twenty-six: your first real machine learning models, evaluated properly.
- Video twenty-seven: turning analysis into a real, effective story.
- Video twenty-eight to thirty: three complete, real end-to-end projects, tying every single skill together.
- That's the whole series: zero to hero, entirely in Python, entirely on real data. Thank you for building all thirty of these with me. If this helped, subscribing and leaving a comment genuinely helps this channel keep making more.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



