Lesson 30 · Python for Retail E-commerce Analytics
Customer Clustering Using K-Means in Python for Retail Analytics
In this lesson you will learn how to segment retail customers using K-Means clustering. Customer segmentation helps businesses group their customers by…
- CoursePython for Retail E-commerce Analytics
- Lesson30 of 43
- Video21 min
- FormatJupyter notebook · 20 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbCustomer Clustering Using K-Means#
- In this lesson you will learn how to segment retail customers using K-Means clustering.
- Customer segmentation helps businesses group their customers by purchasing behavior.
- With clustering, companies can design personalized marketing plans and identify high-value clients.
- You will analyze real retail transaction data to group customers by their spending patterns and frequency.
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
from sklearn.cluster import KMeans
from sklearn.preprocessing import StandardScaler
import warnings
warnings.filterwarnings('ignore')
Core Concepts: Retail Data & Sales Metrics#
- Retail datasets usually include transactions, orders, products, and customer details.
- Important sales metrics include revenue (Quantity x Price), order frequency, and recency.
- Revenue is often miscalculated by forgetting to multiply price and quantity.
- Grouping by the wrong column, such as by product instead of customer, is common among beginners.
- Maintaining correct data types and formatting is key to accurate analytics.
# Load a real online retail transaction dataset
url = 'https://archive.ics.uci.edu/ml/machine-learning-databases/00502/online_retail_II.xlsx'
df = pd.read_excel(url, sheet_name='Year 2010-2011')
df['InvoiceDate'] = pd.to_datetime(df['InvoiceDate'])
print(df.shape)
print(df.head(3))
# Check for missing values in Customer ID and Price/Quantity
missing_customers = df['Customer ID'].isna().sum()
missing_quantities = df['Quantity'].isna().sum()
missing_prices = df['Price'].isna().sum()
print('Missing Customer IDs:', missing_customers)
print('Missing Quantities:', missing_quantities)
print('Missing Prices:', missing_prices)
# Remove transactions with missing Customer ID or negative Quantity/Price
df = df.dropna(subset=['Customer ID'])
df = df[(df['Quantity'] > 0) & (df['Price'] > 0)]
print('After cleaning:', df.shape)
# Calculate the total revenue per transaction
df['Revenue'] = df['Quantity'] * df['Price']
print(df[['Customer ID','Revenue']].head(3))
# Example 1: Compute total revenue for each customer
customer_revenue = df.groupby('Customer ID')['Revenue'].sum().reset_index()
print(customer_revenue.head(3))
# Example 2: Compute number of purchases per customer
customer_freq = df.groupby('Customer ID')['Invoice'].nunique().reset_index(name='NumPurchases')
print(customer_freq.head(3))
# Example 3: Compute recency for each customer (days since last purchase)
latest_date = df['InvoiceDate'].max()
recency = df.groupby('Customer ID')['InvoiceDate'].max().reset_index()
recency['RecencyDays'] = (latest_date - recency['InvoiceDate']).dt.days
recency = recency[['Customer ID','RecencyDays']]
print(recency.head(3))
# Example 4 (intermediate): Merge revenue, frequency, and recency into one dataset
rfm = customer_revenue.merge(customer_freq, on='Customer ID')
rfm = rfm.merge(recency, on='Customer ID')
print(rfm.head(3))
# Example 5 (intermediate): Visualize distribution of RFM features
fig, axes = plt.subplots(1, 3, figsize=(15,4))
axes[0].hist(rfm['Revenue'], bins=30, color='skyblue')
axes[0].set_title('Customer Revenue Distribution')
axes[1].hist(rfm['NumPurchases'], bins=30, color='salmon')
axes[1].set_title('Purchase Frequency')
axes[2].hist(rfm['RecencyDays'], bins=30, color='lightgreen')
axes[2].set_title('Recency (Days)')
plt.tight_layout()
plt.show()
# Example 6 (intermediate): Scale RFM features for K-Means clustering
scaler = StandardScaler()
rfm_scaled = scaler.fit_transform(rfm[['Revenue','NumPurchases','RecencyDays']])
print('Means after scaling:', rfm_scaled.mean(axis=0))
print('Stds after scaling:', rfm_scaled.std(axis=0))
# Example 7 (intermediate): Elbow method to choose number of clusters
wcss = []
for i in range(1, 11):
kmeans = KMeans(n_clusters=i, random_state=42)
kmeans.fit(rfm_scaled)
wcss.append(kmeans.inertia_)
plt.plot(range(1,11), wcss, marker='o')
plt.xlabel('Number of clusters')
plt.ylabel('WCSS (Within-Cluster Sum of Squares)')
plt.title('Elbow Method for Finding Optimal K')
plt.show()
# Example 8 (advanced): Run K-Means with the chosen number of clusters (e.g., 4)
k = 4
kmeans = KMeans(n_clusters=k, random_state=42)
clusters = kmeans.fit_predict(rfm_scaled)
rfm['Cluster'] = clusters
print(rfm.head(3))
# Example 9 (advanced): Profile customer clusters
cluster_summary = rfm.groupby('Cluster').agg({
'Revenue': 'mean',
'NumPurchases': 'mean',
'RecencyDays': 'mean',
'Customer ID': 'count'
}).rename(columns={'Customer ID': 'NumCustomers'}).reset_index()
print(cluster_summary)
# Example 10 (advanced): Visualize clusters in 2D (using two principal features)
plt.figure(figsize=(8,6))
plt.scatter(rfm['Revenue'], rfm['RecencyDays'], c=rfm['Cluster'], cmap='tab10', alpha=0.7)
plt.xlabel('Revenue')
plt.ylabel('RecencyDays')
plt.title('Customer Clusters by Revenue and Recency')
plt.colorbar(label='Cluster')
plt.show()
# Error handling: What if a groupby fails due to missing columns?
try:
df.groupby('NonExistentColumn').sum()
except Exception as e:
print('Error:', e)
# Error handling: Incorrect aggregation can break customer analysis
try:
df['Total'] = df['Quantity'] + df['Price'] # Incorrect: adding, not multiplying
print(df[['Quantity','Price','Total']].head(3))
except Exception as e:
print('Error:', e)
# Error handling: Grouping by Product instead of Customer
product_agg = df.groupby('Description')['Revenue'].sum().reset_index()
print(product_agg.head(3))
# Best practices: Always validate cluster quality by profiling each cluster
for cluster in cluster_summary['Cluster']:
print(f"Cluster {cluster}: Avg Revenue ${cluster_summary.loc[cluster,'Revenue']:.2f}, "
f"Avg Purchases {cluster_summary.loc[cluster,'NumPurchases']:.1f}, "
f"Recency {cluster_summary.loc[cluster,'RecencyDays']:.1f} days, "
f"Customers: {cluster_summary.loc[cluster,'NumCustomers']}")
Best Practices: Customer Segmentation Patterns#
- Use clear and interpretable features: revenue, frequency, and recency are a solid foundation.
- Standardize all feature values before clustering (especially when units differ).
- Visualize both distributions and cluster assignments for actionable insights.
- Profile and describe each cluster to discover what makes each group unique.
# End-to-end retail analytics problem: Identify and describe "VIP" customers
# Definition: VIPs = customers in the cluster with highest average revenue
vip_cluster = cluster_summary.loc[cluster_summary['Revenue'].idxmax(), 'Cluster']
vip_customers = rfm[rfm['Cluster'] == vip_cluster]
print(f"VIP customer count: {len(vip_customers)}")
print(vip_customers[['Customer ID', 'Revenue', 'NumPurchases', 'RecencyDays']].head())
Congratulations! Keep Practicing Advanced Retail Analytics#
- Try different feature combinations or clustering algorithms to segment your customers.
- For further learning, search YouTube for "Retail Customer Segmentation Python".
- Share your findings and cluster plots with your study group or supervisor.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



