Lesson 28 · Data Mining
Understanding Socio-Economic Segmentation Through Hands-On Clustering Techniques
Welcome to your hands-on clustering project! This week we will use the Mall Customers Dataset. You will learn how to explore, clean, and segment customers…
- CourseData Mining
- Lesson28 of 31
- Video16 min
- FormatJupyter notebook · 15 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbWeek 7 Hands-on Clustering Project: Socio-Economic Segmentation#
Welcome to your hands-on clustering project! This week we will use the Mall Customers Dataset.
You will learn how to explore, clean, and segment customers using clustering techniques.
By the end, you will be able to find groups of similar shoppers a skill data scientists use for better marketing, product design, and more.
Let us get started!
# Suppress all warnings for a cleaner experience
import warnings; warnings.filterwarnings("ignore")
import numpy as np
np.random.seed(42)
Data setup (Mall Customers Dataset)#
We will load the Mall Customers Dataset. This dataset contains info like:
- CustomerID
- Gender
- Age
- Annual Income (k$)
- Spending Score (1-100)
Let us see what the data looks like.
import pandas as pd
url = 'https://gist.githubusercontent.com/pravalliyaram/5c05f43d2351249927b8a3f3cc3e5ecf/raw/Mall_Customers.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
Step 1: Exploring the Data#
Let us quickly see what columns we have and get some summary stats.
This helps us spot missing or odd values early.
print(df.columns)
print(df.info())
print(df.describe())
Step 2: Checking for Missing Values#
Missing data can be a major problem in analysis.
We will check if any values are missing.
print(df.isnull().sum())
# For now, let us drop the 'CustomerID' column since it does not help with clustering
df = df.drop('CustomerID', axis=1)
Step 3: Visualizing Age and Income#
Let us visualize the Age and Annual Income columns.
This helps us understand the distribution and spot outliers.
import matplotlib.pyplot as plt
plt.figure(figsize=(10,4))
plt.subplot(1,2,1); plt.hist(df['Age'], bins=15, color='skyblue', edgecolor='k'); plt.title('Age Distribution'); plt.xlabel('Age'); plt.ylabel('Count')
plt.subplot(1,2,2); plt.hist(df['Annual Income (k$)'], bins=15, color='salmon', edgecolor='k'); plt.title('Income Distribution'); plt.xlabel('Annual Income (k$)'); plt.ylabel('Count')
plt.tight_layout(); plt.show()
# Let us check the number of male and female customers
print(df['Gender'].value_counts())
# For clustering, we should convert 'Gender' to numbers: 0 for Female, 1 for Male
df['Gender'] = df['Gender'].map({'Female': 0, 'Male': 1})
Step 4: Feature Scaling#
Features with bigger values can impact clustering more.
We will scale features so all have equal weight.
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
scaled_features = scaler.fit_transform(df)
print(scaled_features[:3])
Step 5: K-Means Clustering#
Now for the exciting part: Clustering!
K-Means helps group similar customers. You choose the number of groups (called clusters).
Let us try clustering into 3 groups.
from sklearn.cluster import KMeans
kmeans = KMeans(n_clusters=3, random_state=42)
clusters = kmeans.fit_predict(scaled_features)
print('Cluster labels:', clusters[:10])
# Let us add our cluster results back to the dataframe for easy use
df['Cluster'] = clusters
# Look at a few customers from each group
for c in range(3):
print(f'Cluster {c}\n', df[df['Cluster']==c].head(2), '\n')
Step 6: Visualizing the Clusters#
A plot helps us see the customer segments in action.
Let us plot clusters using Annual Income and Spending Score.
colors = ['red', 'green', 'blue']
plt.figure(figsize=(8,5))
for c in range(3):
plt.scatter(df[df['Cluster']==c]['Annual Income (k$)'], df[df['Cluster']==c]['Spending Score (1-100)'],
color=colors[c], label=f'Cluster {c}')
plt.xlabel('Annual Income (k$)')
plt.ylabel('Spending Score (1-100)')
plt.title('Mall Customers Clusters')
plt.legend()
plt.show()
# Optional: How many clusters would be best? Let us use the elbow method to find out
inertia = []
for k in range(1,8):
model = KMeans(n_clusters=k, random_state=42).fit(scaled_features)
inertia.append(model.inertia_)
plt.plot(range(1,8), inertia, marker='o'); plt.xlabel('Number of clusters'); plt.ylabel('Inertia'); plt.title('Elbow Plot')
plt.show()
Step 7: Customer Segment Analysis#
Can you describe what kind of customers are in each segment?
Let us see the average features for each group.
print(df.groupby('Cluster').mean())
Challenge Exercise#
Try these challenges:
- Cluster with 4 different groups. What changes?
- Find the average age for each new group.
- Create a scatter plot using Age instead of Annual Income. Are the clusters clear?
Share your thoughts below!
Recap: What You Learned#
You:
- Loaded and inspected real customer data
- Prepared it for analysis
- Used K-Means for clustering
- Explored results with plots and group averages
- Practiced describing each segment
Great job! See you next time.
Want More Data Mining Fun?#
Hit subscribe and turn on notifications for more videos.
Let us keep learning together!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



