Mathew K Analytics

Lesson 28 · Data Mining

Understanding Socio-Economic Segmentation Through Hands-On Clustering Techniques

Welcome to your hands-on clustering project! This week we will use the Mall Customers Dataset. You will learn how to explore, clean, and segment customers…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Week 7 Hands-on Clustering Project: Socio-Economic Segmentation#

Welcome to your hands-on clustering project! This week we will use the Mall Customers Dataset.

You will learn how to explore, clean, and segment customers using clustering techniques.

By the end, you will be able to find groups of similar shoppers a skill data scientists use for better marketing, product design, and more.

Let us get started!

# Suppress all warnings for a cleaner experience
import warnings; warnings.filterwarnings("ignore")
import numpy as np
np.random.seed(42)

Data setup (Mall Customers Dataset)#

We will load the Mall Customers Dataset. This dataset contains info like:

  • CustomerID
  • Gender
  • Age
  • Annual Income (k$)
  • Spending Score (1-100)

Let us see what the data looks like.

import pandas as pd
url = 'https://gist.githubusercontent.com/pravalliyaram/5c05f43d2351249927b8a3f3cc3e5ecf/raw/Mall_Customers.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
(200, 5)
   CustomerID  Gender  Age  Annual Income (k$)  Spending Score (1-100)
0           1    Male   19                  15                      39
1           2    Male   21                  15                      81
2           3  Female   20                  16                       6

Step 1: Exploring the Data#

Let us quickly see what columns we have and get some summary stats.

This helps us spot missing or odd values early.

print(df.columns)
print(df.info())
print(df.describe())
Index(['CustomerID', 'Gender', 'Age', 'Annual Income (k$)',
       'Spending Score (1-100)'],
      dtype='object')
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 200 entries, 0 to 199
Data columns (total 5 columns):
 #   Column                  Non-Null Count  Dtype 
---  ------                  --------------  ----- 
 0   CustomerID              200 non-null    int64 
 1   Gender                  200 non-null    object
 2   Age                     200 non-null    int64 
 3   Annual Income (k$)      200 non-null    int64 
 4   Spending Score (1-100)  200 non-null    int64 
dtypes: int64(4), object(1)
memory usage: 7.9+ KB
None
       CustomerID         Age  Annual Income (k$)  Spending Score (1-100)
count  200.000000  200.000000          200.000000              200.000000
mean   100.500000   38.850000           60.560000               50.200000
std     57.879185   13.969007           26.264721               25.823522
min      1.000000   18.000000           15.000000                1.000000
25%     50.750000   28.750000           41.500000               34.750000
50%    100.500000   36.000000           61.500000               50.000000
75%    150.250000   49.000000           78.000000               73.000000
max    200.000000   70.000000          137.000000               99.000000

Step 2: Checking for Missing Values#

Missing data can be a major problem in analysis.

We will check if any values are missing.

print(df.isnull().sum())
CustomerID                0
Gender                    0
Age                       0
Annual Income (k$)        0
Spending Score (1-100)    0
dtype: int64
# For now, let us drop the 'CustomerID' column since it does not help with clustering
df = df.drop('CustomerID', axis=1)

Step 3: Visualizing Age and Income#

Let us visualize the Age and Annual Income columns.

This helps us understand the distribution and spot outliers.

import matplotlib.pyplot as plt
plt.figure(figsize=(10,4))
plt.subplot(1,2,1); plt.hist(df['Age'], bins=15, color='skyblue', edgecolor='k'); plt.title('Age Distribution'); plt.xlabel('Age'); plt.ylabel('Count')
plt.subplot(1,2,2); plt.hist(df['Annual Income (k$)'], bins=15, color='salmon', edgecolor='k'); plt.title('Income Distribution'); plt.xlabel('Annual Income (k$)'); plt.ylabel('Count')
plt.tight_layout(); plt.show()
No description has been provided for this image
# Let us check the number of male and female customers
print(df['Gender'].value_counts())
Gender
Female    112
Male       88
Name: count, dtype: int64
# For clustering, we should convert 'Gender' to numbers: 0 for Female, 1 for Male
df['Gender'] = df['Gender'].map({'Female': 0, 'Male': 1})

Step 4: Feature Scaling#

Features with bigger values can impact clustering more.

We will scale features so all have equal weight.

from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
scaled_features = scaler.fit_transform(df)
print(scaled_features[:3])
[[ 1.12815215 -1.42456879 -1.73899919 -0.43480148]
 [ 1.12815215 -1.28103541 -1.73899919  1.19570407]
 [-0.88640526 -1.3528021  -1.70082976 -1.71591298]]

Step 5: K-Means Clustering#

Now for the exciting part: Clustering!

K-Means helps group similar customers. You choose the number of groups (called clusters).

Let us try clustering into 3 groups.

from sklearn.cluster import KMeans
kmeans = KMeans(n_clusters=3, random_state=42)
clusters = kmeans.fit_predict(scaled_features)
print('Cluster labels:', clusters[:10])
Cluster labels: [2 2 0 2 0 2 0 2 0 2]
# Let us add our cluster results back to the dataframe for easy use
df['Cluster'] = clusters
# Look at a few customers from each group
for c in range(3):
    print(f'Cluster {c}\n', df[df['Cluster']==c].head(2), '\n')
    
Cluster 0
    Gender  Age  Annual Income (k$)  Spending Score (1-100)  Cluster
2       0   20                  16                       6        0
4       0   31                  17                      40        0 

Cluster 1
      Gender  Age  Annual Income (k$)  Spending Score (1-100)  Cluster
126       1   43                  71                      35        1
128       1   59                  71                      11        1 

Cluster 2
    Gender  Age  Annual Income (k$)  Spending Score (1-100)  Cluster
0       1   19                  15                      39        2
1       1   21                  15                      81        2 

Step 6: Visualizing the Clusters#

A plot helps us see the customer segments in action.

Let us plot clusters using Annual Income and Spending Score.

colors = ['red', 'green', 'blue']
plt.figure(figsize=(8,5))
for c in range(3):
    plt.scatter(df[df['Cluster']==c]['Annual Income (k$)'], df[df['Cluster']==c]['Spending Score (1-100)'],
                color=colors[c], label=f'Cluster {c}')
plt.xlabel('Annual Income (k$)')
plt.ylabel('Spending Score (1-100)')
plt.title('Mall Customers Clusters')
plt.legend()
plt.show()
No description has been provided for this image
# Optional: How many clusters would be best? Let us use the elbow method to find out
inertia = []
for k in range(1,8):
    model = KMeans(n_clusters=k, random_state=42).fit(scaled_features)
    inertia.append(model.inertia_)
plt.plot(range(1,8), inertia, marker='o'); plt.xlabel('Number of clusters'); plt.ylabel('Inertia'); plt.title('Elbow Plot')
plt.show()
No description has been provided for this image

Step 7: Customer Segment Analysis#

Can you describe what kind of customers are in each segment?

Let us see the average features for each group.

print(df.groupby('Cluster').mean())
           Gender        Age  Annual Income (k$)  Spending Score (1-100)
Cluster                                                                 
0        0.394366  52.169014           46.676056               39.295775
1        0.628571  40.228571           91.342857               20.628571
2        0.404255  28.276596           59.585106               69.446809

Challenge Exercise#

Try these challenges:

  1. Cluster with 4 different groups. What changes?
  2. Find the average age for each new group.
  3. Create a scatter plot using Age instead of Annual Income. Are the clusters clear?

Share your thoughts below!

Recap: What You Learned#

You:

  • Loaded and inspected real customer data
  • Prepared it for analysis
  • Used K-Means for clustering
  • Explored results with plots and group averages
  • Practiced describing each segment

Great job! See you next time.

Want More Data Mining Fun?#

Hit subscribe and turn on notifications for more videos.

Let us keep learning together!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.