Lesson 69 · Python Fundamentals
Customer Segmentation Using PCA and KMeans Clustering in Python: A Step-by-Step Guide
Welcome! Today you will learn how to group customers using real data. You will use important tools called PCA and KMeans. These help discover patterns and…
- CoursePython Fundamentals
- Lesson69 of 22
- Video15 min
- FormatJupyter notebook · 18 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynb
Customer Segmentation with PCA and KMeans in Python#
Welcome! Today you will learn how to group customers using real data.
You will use important tools called PCA and KMeans.
These help discover patterns and clusters in messy information.
No experience is needed! Let us explore hands-on.
# Let us start by making sure warnings do not distract us
import warnings
warnings.filterwarnings('ignore')
# Now, import pandas for easy table data
import pandas as pd
What is customer segmentation?#
Companies often want to group customers by habits or traits.
Grouping helps them target special offers and better understand shoppers.
Today, you will use real shopper data to see clusters and patterns.
# Data setup
url = "https://raw.githubusercontent.com/mwaskom/seaborn-data/master/mall_customers.csv"
df = pd.read_csv(url)
print("Data shape:", df.shape)
df.head()
Exploring the dataset#
Understanding the info inside your data is a huge first step.
Check out the features:
- CustomerID: Unique number for each shopper
- Genre: Male or Female
- Age: Years old
- Annual Income (k$): Yearly earnings, in thousands
- Spending Score (1-100): How much they spend
Let us look a bit deeper.
# Describe gives us basic stats for numbers
df.describe()
# Let us peek at unique genres
print(df['Genre'].value_counts())
Getting data ready for analysis#
Some analysis tools only understand numbers, not words.
So, we must turn 'Genre' into a number first.
This is called encoding.
# Convert Genre: Female -> 0, Male -> 1
df['Gender_Code'] = df['Genre'].map({'Female': 0, 'Male': 1})
df[['Genre', 'Gender_Code']].head()
# Pick only numeric features for our clustering
features = [
'Age',
'Annual Income (k$)',
'Spending Score (1-100)',
'Gender_Code'
]
X = df[features]
What is PCA?#
PCA stands for Principal Component Analysis.
It helps us turn messy features into simple patterns.
PCA finds the big trends and lets us see data in fewer dimensions.
This helps with exploring and clustering.
# Let us run PCA to shrink our features to 2 principal components
from sklearn.preprocessing import StandardScaler
from sklearn.decomposition import PCA
# Scale our data so every column is fair
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)
pca = PCA(n_components=2, random_state=0)
X_pca = pca.fit_transform(X_scaled)
print("PCA-shape:", X_pca.shape)
# Plot the two main PCA axes to see any clusters
import matplotlib.pyplot as plt
plt.figure(figsize=(7,6))
plt.scatter(X_pca[:, 0], X_pca[:, 1], c='skyblue', alpha=0.5)
plt.xlabel('PC1')
plt.ylabel('PC2')
plt.title('Customers after PCA')
plt.show()
Time for clustering! What is KMeans?#
KMeans finds groups (clusters) by guessing where group centers might be.
KMeans keeps moving the centers to fit real patterns.
We will start with 5 clusters, just for practice.
# Group customers into 5 clusters using KMeans
from sklearn.cluster import KMeans
kmeans = KMeans(n_clusters=5, random_state=0)
clusters = kmeans.fit_predict(X_pca)
df['Cluster'] = clusters
df[['CustomerID', 'Cluster']].head()
# Visualize clusters in PCA space
plt.figure(figsize=(8,6))
scatter = plt.scatter(X_pca[:, 0], X_pca[:, 1], c=clusters, cmap='Set2', alpha=0.7)
plt.xlabel('PC1')
plt.ylabel('PC2')
plt.title('Customer Clusters (KMeans, k=5)')
plt.colorbar(scatter, label='Cluster')
plt.show()
Real-world use: What are clusters good for?#
Customer groups can help design special offers.
Managers might focus on high spenders or certain ages.
Not all clusters mean profit always explore with a goal in mind.
# See what an average customer looks like in each cluster
df.groupby('Cluster')[features].mean()
# Let us pick a cluster to analyze up-close
cluster_num = int(input("Enter a cluster number (0-4) to inspect: "))
sample = df[df['Cluster'] == cluster_num].head()
sample
# Mini-project: Build your own segmentation rule
print("Let us make a group for young, high-spending customers.")
mask = (df['Age'] < 30) & (df['Spending Score (1-100)'] > 60)
young_high_spenders = df[mask]
print("There are", len(young_high_spenders), "customers in this group.")
young_high_spenders.head()
# Troubleshooting: What if you get an error?
# Common issues: missing quotes, wrong spelling, wrong column names.
# If an error appears, read the message. It tells you what is broken.
# Try to match the error's words to your code to find the fix.
# Let us try an error. Uncomment the next line to see one.
# print(df['NotAColumn'].head())
# Best practices: Always check for missing data
missing = df.isnull().sum()
print("Missing values per column:")
print(missing)
# Extra tip: Try more or fewer clusters
for k in [2, 3, 6]:
kmeans = KMeans(n_clusters=k, random_state=0)
clusters_try = kmeans.fit_predict(X_pca)
print(f"Number of clusters: {k} - Unique labels: {len(set(clusters_try))}")
Recap and next steps#
You have learned how to load, prepare, and group real customer data.
You tried PCA for simplifying features.
KMeans found pattern groups to help businesses take action.
Keep exploring and practicing new datasets!
Try more on your own!#
Can you:
- Plot age versus spending score for one cluster?
- Try other real datasets like Titanic or Iris?
- Change PCA to 3 components and plot?
If you enjoyed this, subscribe for more Python, and share your clusters below!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



