Mathew K Analytics

Lesson 24 · Data Mining

Understanding the k-Means Clustering Algorithm: A Clear Step-by-Step Guide

In this lesson, we will explore how k-Means clustering works and apply it to real data. We will use the Mall Customers Dataset to find patterns in how…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Week 7: Introduction to k-Means Clustering with the Mall Customers Dataset#

In this lesson, we will explore how k-Means clustering works and apply it to real data.

We will use the Mall Customers Dataset to find patterns in how customers behave.

Clustering is a key task in data mining for revealing groups that are not easily seen just by looking.

What is Clustering?#

Clustering is a way to group similar data points together, based only on their features.

It is a type of unsupervised learning, meaning we do not give the computer labels or answers.

Instead, the computer tries to find its own groups within the data.

import warnings
import numpy as np
np.random.seed(42)
warnings.filterwarnings('ignore')
# We suppress any warnings, so the output stays clean.
# Data setup (Mall Customers Dataset)
import pandas as pd
url = 'https://gist.githubusercontent.com/pravalliyaram/5c05f43d2351249927b8a3f3cc3e5ecf/raw/Mall_Customers.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
(200, 5)
   CustomerID  Gender  Age  Annual Income (k$)  Spending Score (1-100)
0           1    Male   19                  15                      39
1           2    Male   21                  15                      81
2           3  Female   20                  16                       6

Exploring Our Data#

This dataset shows different people who shop at a mall.

Each customer has several features: ID, Gender, Age, Annual Income, and Spending Score.

Spending Score is a value assigned by the mall to rank how much they spend.

Our job will be to find groups of customers who behave alike.

# Check for missing values
print(df.isnull().sum())
CustomerID                0
Gender                    0
Age                       0
Annual Income (k$)        0
Spending Score (1-100)    0
dtype: int64
# Quick summary statistics
print(df.describe())
       CustomerID         Age  Annual Income (k$)  Spending Score (1-100)
count  200.000000  200.000000          200.000000              200.000000
mean   100.500000   38.850000           60.560000               50.200000
std     57.879185   13.969007           26.264721               25.823522
min      1.000000   18.000000           15.000000                1.000000
25%     50.750000   28.750000           41.500000               34.750000
50%    100.500000   36.000000           61.500000               50.000000
75%    150.250000   49.000000           78.000000               73.000000
max    200.000000   70.000000          137.000000               99.000000

Data Preprocessing#

k-Means can only work with numbers.

We will choose only the features we want, and turn them into numbers if needed.

# Let us select features for clustering
X = df[['Age', 'Annual Income (k$)', 'Spending Score (1-100)']]
# Visualizing the input features
import matplotlib.pyplot as plt
plt.scatter(X['Annual Income (k$)'], X['Spending Score (1-100)'], c='blue', alpha=0.5)
plt.xlabel('Annual Income (k$)')
plt.ylabel('Spending Score (1-100)')
plt.title('Customer distribution by Income and Spending Score')
plt.show()
No description has been provided for this image

Scaling Our Data#

Clustering often works better if each feature has the same scale.

We will use a tool to make sure each number is on a similar scale.

from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)

What is k-Means?#

k-Means is an algorithm that puts data into k groups or clusters.

It works by picking k starting points, then moving them to better spots until the groups are clear.

You have to pick how many groups (k) you want to start with.

# Let us try k-Means with 3 clusters
from sklearn.cluster import KMeans
kmeans = KMeans(n_clusters=3, random_state=42)
kmeans.fit(X_scaled)
labels = kmeans.labels_
print('Cluster labels:', labels[:10])
Cluster labels: [2 2 2 2 2 2 0 2 0 2]
# Add cluster label to the original data
df['Cluster'] = labels
print(df[['CustomerID', 'Cluster']].head(10))
   CustomerID  Cluster
0           1        2
1           2        2
2           3        2
3           4        2
4           5        2
5           6        2
6           7        0
7           8        2
8           9        0
9          10        2
# Visualizing clusters
plt.figure(figsize=(7,5))
plt.scatter(X['Annual Income (k$)'], X['Spending Score (1-100)'], c=labels, cmap='viridis', alpha=0.7)
plt.xlabel('Annual Income (k$)')
plt.ylabel('Spending Score (1-100)')
plt.title('k-Means Clusters of Customers')
plt.colorbar(label='Cluster Label')
plt.show()
No description has been provided for this image
# Finding the 'right' k using the elbow method
inertia = []
ks = range(1, 11)
for k in ks:
    km = KMeans(n_clusters=k, random_state=42)
    km.fit(X_scaled)
    inertia.append(km.inertia_)

plt.plot(ks, inertia, marker='o')
plt.xlabel('Number of clusters (k)')
plt.ylabel('Inertia')
plt.title('Elbow Method for Finding k')
plt.show()
No description has been provided for this image
# Characteristics of each cluster
cluster_summary = df.groupby('Cluster')[['Age', 'Annual Income (k$)', 'Spending Score (1-100)']].mean()
print(cluster_summary)
               Age  Annual Income (k$)  Spending Score (1-100)
Cluster                                                       
0        50.406250           60.468750               33.343750
1        32.853659           87.341463               79.975610
2        25.142857           43.269841               56.507937

Practice: Try It Yourself!#

Change the number of clusters to 4 or 5 and run the k-Means steps again.

Which groups can you see in your results?

# Test your understanding: What happens if we skip scaling?
kmeans_no_scale = KMeans(n_clusters=3, random_state=42).fit(X)
labels_ns = kmeans_no_scale.labels_
plt.scatter(X['Annual Income (k$)'], X['Spending Score (1-100)'], c=labels_ns, cmap='plasma', alpha=0.7)
plt.xlabel('Annual Income (k$)')
plt.ylabel('Spending Score (1-100)')
plt.title('k-Means Without Scaling')
plt.colorbar(label='Cluster Label')
plt.show()
No description has been provided for this image

Recap: What We Learned#

Today we learned how to use k-Means clustering on customer data.

We saw how scaling and picking k are both important.

Using k-Means, we found natural customer groups and could start to see who spends differently in the mall.

Next steps#

Try playing with different features or your own datasets.

Share your results and ideas in the comments.

Remember to subscribe for more beginner-friendly data mining tutorials!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.