Mathew K Analytics

Lesson 22 · Python For Machine Learning

Understanding KMeans Clustering in Python: A Guide to Unsupervised Machine Learning

Welcome! In this lesson, you will discover how to group similar data automatically using KMeans clustering. No experience needed you will learn step by…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb
 

Getting Started with KMeans Clustering in Python#

Welcome! In this lesson, you will discover how to group similar data automatically using KMeans clustering.

No experience needed you will learn step by step.

What is Clustering?

Clustering is a way for computers to sort things into groups based on patterns.

Why learn KMeans?

KMeans is a popular and simple clustering algorithm.

You will use Python and real-world data to make clusters.

Let us begin our journey!

# Let us silence warnings for a cleaner notebook
import warnings
warnings.filterwarnings('ignore')
# Let us import some useful tools for our lesson
import pandas as pd
from sklearn.cluster import KMeans
import matplotlib.pyplot as plt
%matplotlib inline

Data setup#

We will use the Mall Customers dataset. It contains information about people shopping at a mall.

You will see how clustering can help find groups, like similar types of shoppers.

# Load the mall customers data
data_url = "https://raw.githubusercontent.com/rithwik00/Mall-Customer-Dataset/main/Mall_Customers.csv"
df = pd.read_csv(data_url)
df.columns = df.columns.str.strip()

# Show the shape and a quick peek at the data
print('Shape of data:', df.shape)
df.head()
Shape of data: (200, 5)
CustomerID Gender Age Annual Income (k$) Spending Score (1-100)
0 1 Male 19 15 39
1 2 Male 21 15 81
2 3 Female 20 16 6
3 4 Female 23 16 77
4 5 Female 31 17 40
# Let us check if our data has missing values
df.isnull().sum()
CustomerID                0
Gender                    0
Age                       0
Annual Income (k$)        0
Spending Score (1-100)    0
dtype: int64

Choosing features for clustering#

We do not need every column for clustering.

Let us pick two columns to group people: 'Annual Income (k$)' and 'Spending Score (1-100)'.

These will help us find patterns in how much people earn and spend.

# Select the columns we will use for clustering
features = df[["Annual Income (k$)", "Spending Score (1-100)"]]
features.head()
Annual Income (k$) Spending Score (1-100)
0 15 39
1 15 81
2 16 6
3 16 77
4 17 40

Visualizing the data#

It is helpful to see data before clustering.

Let us make a plot to spot any groups by eye.

# Plot the customers by income and spending
plt.figure(figsize=(8,6))
plt.scatter(features["Annual Income (k$)"], features["Spending Score (1-100)"])
plt.xlabel("Annual Income (k$)")
plt.ylabel("Spending Score (1-100)")
plt.title("Customers by Income and Spending Score")
plt.show()
No description has been provided for this image

What is KMeans Clustering?#

KMeans is a way to group data into clusters automatically.

You tell KMeans how many clusters to find. The algorithm puts similar points together.

Each group is called a cluster. A "mean" is just the average location in each group.

# Try clustering with 3 groups first
kmeans = KMeans(n_clusters=3, random_state=42)
kmeans.fit(features)
labels = kmeans.labels_
print(labels[:10])
[2 2 2 2 2 2 2 2 2 2]
# Change the number of clusters using input()
k = int(input("How many clusters would you like to try? "))
kmeans_alt = KMeans(n_clusters=k, random_state=42)
kmeans_alt.fit(features)
df["Cluster"] = kmeans_alt.labels_
df[["Annual Income (k$)", "Spending Score (1-100)", "Cluster"]].head()
Annual Income (k$) Spending Score (1-100) Cluster
0 15 39 4
1 15 81 2
2 16 6 4
3 16 77 2
4 17 40 4
 
# Plot the clusters with different colors
plt.figure(figsize=(8,6))
for i in range(k):
    plt.scatter(
        features[kmeans_alt.labels_ == i]["Annual Income (k$)"],
        features[kmeans_alt.labels_ == i]["Spending Score (1-100)"],
        label=f"Cluster {i+1}",
        alpha=0.7
    )
plt.xlabel("Annual Income (k$)")
plt.ylabel("Spending Score (1-100)")
plt.title(f"KMeans with {k} Clusters")
plt.legend()
plt.show()
No description has been provided for this image
# Check where each cluster center is
centers = kmeans_alt.cluster_centers_
print("Cluster Centers:")
print(centers)
Cluster Centers:
[[55.2962963  49.51851852]
 [86.53846154 82.12820513]
 [25.72727273 79.36363636]
 [88.2        17.11428571]
 [26.30434783 20.91304348]]
# What happens if we pick way too many clusters?
k_too_many = 12
kmeans_many = KMeans(n_clusters=k_too_many, random_state=42)
kmeans_many.fit(features)
plt.figure(figsize=(8,6))
for i in range(k_too_many):
    plt.scatter(
        features[kmeans_many.labels_ == i]["Annual Income (k$)"],
        features[kmeans_many.labels_ == i]["Spending Score (1-100)"],
        label=f"Cluster {i+1}",
        alpha=0.7
    )
plt.xlabel("Annual Income (k$)")
plt.ylabel("Spending Score (1-100)")
plt.title("KMeans with Too Many Clusters")
plt.legend(bbox_to_anchor=(1.05, 1), loc='upper left')
plt.show()
No description has been provided for this image
# How to choose the best number of clusters? The Elbow method.
inertias = []
for k_try in range(1, 11):
    model = KMeans(n_clusters=k_try, random_state=42)
    model.fit(features)
    inertias.append(model.inertia_)
plt.figure(figsize=(6,4))
plt.plot(range(1, 11), inertias, marker='o')
plt.xlabel("Number of Clusters")
plt.ylabel("Inertia (sum of distances)")
plt.title("Elbow Method for Choosing k")
plt.show()
No description has been provided for this image
# Add clusters to our table, and look at each group
best_k = 5
final_kmeans = KMeans(n_clusters=best_k, random_state=42)
df["Cluster"] = final_kmeans.fit_predict(features)
df.groupby("Cluster").mean(numeric_only=True)
CustomerID Age Annual Income (k$) Spending Score (1-100)
Cluster
0 86.320988 42.716049 55.296296 49.518519
1 162.000000 32.692308 86.538462 82.128205
2 23.090909 25.272727 25.727273 79.363636
3 164.371429 41.114286 88.200000 17.114286
4 23.000000 45.217391 26.304348 20.913043
# Add some new data and predict which cluster it would join
new_customer = [[80, 50]]
cluster_for_new = final_kmeans.predict(new_customer)
print(f"Cluster for income 80k$ and spending score 50 is: {cluster_for_new[0]}")
Cluster for income 80k$ and spending score 50 is: 0
# Remove the Cluster column if we are finished analyzing
df = df.drop(columns=["Cluster"])
df.head()
CustomerID Gender Age Annual Income (k$) Spending Score (1-100)
0 1 Male 19 15 39
1 2 Male 21 15 81
2 3 Female 20 16 6
3 4 Female 23 16 77
4 5 Female 31 17 40
# How does KMeans handle out-of-range values?
wild_customer = [[1000, 0]]
wild_group = final_kmeans.predict(wild_customer)
print(f"Extreme customer goes to cluster: {wild_group[0]}")
Extreme customer goes to cluster: 3

Troubleshooting: Common errors#

  • KMeans will not accept missing values.

  • You need only numbers to clustertext will not work.

  • Always pick a random_state for reproducible results.

If you get confused, check these spots first.

# Extra tip: Resetting your notebook
df = pd.read_csv(data_url)
df.columns = df.columns.str.strip()
features = df[["Annual Income (k$)", "Spending Score (1-100)"]]
# Challenge: Try clustering by age and income instead
features_challenge = df[["Age", "Annual Income (k$)"]]
challenge_kmeans = KMeans(n_clusters=4, random_state=42)
challenge_kmeans.fit(features_challenge)
labels_challenge = challenge_kmeans.labels_
plt.figure(figsize=(8,6))
for i in range(4):
    plt.scatter(
        features_challenge[labels_challenge == i]["Age"],
        features_challenge[labels_challenge == i]["Annual Income (k$)"],
        label=f"Cluster {i+1}",
        alpha=0.7
    )
plt.xlabel("Age")
plt.ylabel("Annual Income (k$)")
plt.title("Clusters by Age and Income")
plt.legend()
plt.show()
No description has been provided for this image

Recap: What have you learned?#

  • What clustering is and how to use KMeans
  • How to load real-world data and prepare it
  • How to pick the right features and cluster number
  • How to plot and understand the groups
  • Tips to avoid common mistakes
  • How to try new features for practice

Great job! Practice more to become confident.

Thank you for learning KMeans with me!#

If you enjoyed this lesson, please like, subscribe, and share.

Try clustering on another dataset, and let us know your results in the comments!

See you in the next Python adventure!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.