Lesson 24 · Data Mining
Understanding the k-Means Clustering Algorithm: A Clear Step-by-Step Guide
In this lesson, we will explore how k-Means clustering works and apply it to real data. We will use the Mall Customers Dataset to find patterns in how…
- CourseData Mining
- Lesson24 of 31
- Video16 min
- FormatJupyter notebook · 13 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbWeek 7: Introduction to k-Means Clustering with the Mall Customers Dataset#
In this lesson, we will explore how k-Means clustering works and apply it to real data.
We will use the Mall Customers Dataset to find patterns in how customers behave.
Clustering is a key task in data mining for revealing groups that are not easily seen just by looking.
What is Clustering?#
Clustering is a way to group similar data points together, based only on their features.
It is a type of unsupervised learning, meaning we do not give the computer labels or answers.
Instead, the computer tries to find its own groups within the data.
import warnings
import numpy as np
np.random.seed(42)
warnings.filterwarnings('ignore')
# We suppress any warnings, so the output stays clean.
# Data setup (Mall Customers Dataset)
import pandas as pd
url = 'https://gist.githubusercontent.com/pravalliyaram/5c05f43d2351249927b8a3f3cc3e5ecf/raw/Mall_Customers.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
Exploring Our Data#
This dataset shows different people who shop at a mall.
Each customer has several features: ID, Gender, Age, Annual Income, and Spending Score.
Spending Score is a value assigned by the mall to rank how much they spend.
Our job will be to find groups of customers who behave alike.
# Check for missing values
print(df.isnull().sum())
# Quick summary statistics
print(df.describe())
Data Preprocessing#
k-Means can only work with numbers.
We will choose only the features we want, and turn them into numbers if needed.
# Let us select features for clustering
X = df[['Age', 'Annual Income (k$)', 'Spending Score (1-100)']]
# Visualizing the input features
import matplotlib.pyplot as plt
plt.scatter(X['Annual Income (k$)'], X['Spending Score (1-100)'], c='blue', alpha=0.5)
plt.xlabel('Annual Income (k$)')
plt.ylabel('Spending Score (1-100)')
plt.title('Customer distribution by Income and Spending Score')
plt.show()
Scaling Our Data#
Clustering often works better if each feature has the same scale.
We will use a tool to make sure each number is on a similar scale.
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)
What is k-Means?#
k-Means is an algorithm that puts data into k groups or clusters.
It works by picking k starting points, then moving them to better spots until the groups are clear.
You have to pick how many groups (k) you want to start with.
# Let us try k-Means with 3 clusters
from sklearn.cluster import KMeans
kmeans = KMeans(n_clusters=3, random_state=42)
kmeans.fit(X_scaled)
labels = kmeans.labels_
print('Cluster labels:', labels[:10])
# Add cluster label to the original data
df['Cluster'] = labels
print(df[['CustomerID', 'Cluster']].head(10))
# Visualizing clusters
plt.figure(figsize=(7,5))
plt.scatter(X['Annual Income (k$)'], X['Spending Score (1-100)'], c=labels, cmap='viridis', alpha=0.7)
plt.xlabel('Annual Income (k$)')
plt.ylabel('Spending Score (1-100)')
plt.title('k-Means Clusters of Customers')
plt.colorbar(label='Cluster Label')
plt.show()
# Finding the 'right' k using the elbow method
inertia = []
ks = range(1, 11)
for k in ks:
km = KMeans(n_clusters=k, random_state=42)
km.fit(X_scaled)
inertia.append(km.inertia_)
plt.plot(ks, inertia, marker='o')
plt.xlabel('Number of clusters (k)')
plt.ylabel('Inertia')
plt.title('Elbow Method for Finding k')
plt.show()
# Characteristics of each cluster
cluster_summary = df.groupby('Cluster')[['Age', 'Annual Income (k$)', 'Spending Score (1-100)']].mean()
print(cluster_summary)
Practice: Try It Yourself!#
Change the number of clusters to 4 or 5 and run the k-Means steps again.
Which groups can you see in your results?
# Test your understanding: What happens if we skip scaling?
kmeans_no_scale = KMeans(n_clusters=3, random_state=42).fit(X)
labels_ns = kmeans_no_scale.labels_
plt.scatter(X['Annual Income (k$)'], X['Spending Score (1-100)'], c=labels_ns, cmap='plasma', alpha=0.7)
plt.xlabel('Annual Income (k$)')
plt.ylabel('Spending Score (1-100)')
plt.title('k-Means Without Scaling')
plt.colorbar(label='Cluster Label')
plt.show()
Recap: What We Learned#
Today we learned how to use k-Means clustering on customer data.
We saw how scaling and picking k are both important.
Using k-Means, we found natural customer groups and could start to see who spends differently in the mall.
Next steps#
Try playing with different features or your own datasets.
Share your results and ideas in the comments.
Remember to subscribe for more beginner-friendly data mining tutorials!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



