Lesson 22 · Python For Machine Learning
Understanding KMeans Clustering in Python: A Guide to Unsupervised Machine Learning
Welcome! In this lesson, you will discover how to group similar data automatically using KMeans clustering. No experience needed you will learn step by…
- CoursePython For Machine Learning
- Lesson22 of 16
- Video15 min
- FormatJupyter notebook · 20 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynb
Getting Started with KMeans Clustering in Python#
Welcome! In this lesson, you will discover how to group similar data automatically using KMeans clustering.
No experience needed you will learn step by step.
What is Clustering?
Clustering is a way for computers to sort things into groups based on patterns.
Why learn KMeans?
KMeans is a popular and simple clustering algorithm.
You will use Python and real-world data to make clusters.
Let us begin our journey!
# Let us silence warnings for a cleaner notebook
import warnings
warnings.filterwarnings('ignore')
# Let us import some useful tools for our lesson
import pandas as pd
from sklearn.cluster import KMeans
import matplotlib.pyplot as plt
%matplotlib inline
Data setup#
We will use the Mall Customers dataset. It contains information about people shopping at a mall.
You will see how clustering can help find groups, like similar types of shoppers.
# Load the mall customers data
data_url = "https://raw.githubusercontent.com/rithwik00/Mall-Customer-Dataset/main/Mall_Customers.csv"
df = pd.read_csv(data_url)
df.columns = df.columns.str.strip()
# Show the shape and a quick peek at the data
print('Shape of data:', df.shape)
df.head()
# Let us check if our data has missing values
df.isnull().sum()
Choosing features for clustering#
We do not need every column for clustering.
Let us pick two columns to group people: 'Annual Income (k$)' and 'Spending Score (1-100)'.
These will help us find patterns in how much people earn and spend.
# Select the columns we will use for clustering
features = df[["Annual Income (k$)", "Spending Score (1-100)"]]
features.head()
Visualizing the data#
It is helpful to see data before clustering.
Let us make a plot to spot any groups by eye.
# Plot the customers by income and spending
plt.figure(figsize=(8,6))
plt.scatter(features["Annual Income (k$)"], features["Spending Score (1-100)"])
plt.xlabel("Annual Income (k$)")
plt.ylabel("Spending Score (1-100)")
plt.title("Customers by Income and Spending Score")
plt.show()
What is KMeans Clustering?#
KMeans is a way to group data into clusters automatically.
You tell KMeans how many clusters to find. The algorithm puts similar points together.
Each group is called a cluster. A "mean" is just the average location in each group.
# Try clustering with 3 groups first
kmeans = KMeans(n_clusters=3, random_state=42)
kmeans.fit(features)
labels = kmeans.labels_
print(labels[:10])
# Change the number of clusters using input()
k = int(input("How many clusters would you like to try? "))
kmeans_alt = KMeans(n_clusters=k, random_state=42)
kmeans_alt.fit(features)
df["Cluster"] = kmeans_alt.labels_
df[["Annual Income (k$)", "Spending Score (1-100)", "Cluster"]].head()
# Plot the clusters with different colors
plt.figure(figsize=(8,6))
for i in range(k):
plt.scatter(
features[kmeans_alt.labels_ == i]["Annual Income (k$)"],
features[kmeans_alt.labels_ == i]["Spending Score (1-100)"],
label=f"Cluster {i+1}",
alpha=0.7
)
plt.xlabel("Annual Income (k$)")
plt.ylabel("Spending Score (1-100)")
plt.title(f"KMeans with {k} Clusters")
plt.legend()
plt.show()
# Check where each cluster center is
centers = kmeans_alt.cluster_centers_
print("Cluster Centers:")
print(centers)
# What happens if we pick way too many clusters?
k_too_many = 12
kmeans_many = KMeans(n_clusters=k_too_many, random_state=42)
kmeans_many.fit(features)
plt.figure(figsize=(8,6))
for i in range(k_too_many):
plt.scatter(
features[kmeans_many.labels_ == i]["Annual Income (k$)"],
features[kmeans_many.labels_ == i]["Spending Score (1-100)"],
label=f"Cluster {i+1}",
alpha=0.7
)
plt.xlabel("Annual Income (k$)")
plt.ylabel("Spending Score (1-100)")
plt.title("KMeans with Too Many Clusters")
plt.legend(bbox_to_anchor=(1.05, 1), loc='upper left')
plt.show()
# How to choose the best number of clusters? The Elbow method.
inertias = []
for k_try in range(1, 11):
model = KMeans(n_clusters=k_try, random_state=42)
model.fit(features)
inertias.append(model.inertia_)
plt.figure(figsize=(6,4))
plt.plot(range(1, 11), inertias, marker='o')
plt.xlabel("Number of Clusters")
plt.ylabel("Inertia (sum of distances)")
plt.title("Elbow Method for Choosing k")
plt.show()
# Add clusters to our table, and look at each group
best_k = 5
final_kmeans = KMeans(n_clusters=best_k, random_state=42)
df["Cluster"] = final_kmeans.fit_predict(features)
df.groupby("Cluster").mean(numeric_only=True)
# Add some new data and predict which cluster it would join
new_customer = [[80, 50]]
cluster_for_new = final_kmeans.predict(new_customer)
print(f"Cluster for income 80k$ and spending score 50 is: {cluster_for_new[0]}")
# Remove the Cluster column if we are finished analyzing
df = df.drop(columns=["Cluster"])
df.head()
# How does KMeans handle out-of-range values?
wild_customer = [[1000, 0]]
wild_group = final_kmeans.predict(wild_customer)
print(f"Extreme customer goes to cluster: {wild_group[0]}")
Troubleshooting: Common errors#
KMeans will not accept missing values.
You need only numbers to clustertext will not work.
Always pick a random_state for reproducible results.
If you get confused, check these spots first.
# Extra tip: Resetting your notebook
df = pd.read_csv(data_url)
df.columns = df.columns.str.strip()
features = df[["Annual Income (k$)", "Spending Score (1-100)"]]
# Challenge: Try clustering by age and income instead
features_challenge = df[["Age", "Annual Income (k$)"]]
challenge_kmeans = KMeans(n_clusters=4, random_state=42)
challenge_kmeans.fit(features_challenge)
labels_challenge = challenge_kmeans.labels_
plt.figure(figsize=(8,6))
for i in range(4):
plt.scatter(
features_challenge[labels_challenge == i]["Age"],
features_challenge[labels_challenge == i]["Annual Income (k$)"],
label=f"Cluster {i+1}",
alpha=0.7
)
plt.xlabel("Age")
plt.ylabel("Annual Income (k$)")
plt.title("Clusters by Age and Income")
plt.legend()
plt.show()
Recap: What have you learned?#
- What clustering is and how to use KMeans
- How to load real-world data and prepare it
- How to pick the right features and cluster number
- How to plot and understand the groups
- Tips to avoid common mistakes
- How to try new features for practice
Great job! Practice more to become confident.
Thank you for learning KMeans with me!#
If you enjoyed this lesson, please like, subscribe, and share.
Try clustering on another dataset, and let us know your results in the comments!
See you in the next Python adventure!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



