Mathew K Analytics

Lesson 30 · Python for Retail E-commerce Analytics

Customer Clustering Using K-Means in Python for Retail Analytics

In this lesson you will learn how to segment retail customers using K-Means clustering. Customer segmentation helps businesses group their customers by…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Customer Clustering Using K-Means#

  • In this lesson you will learn how to segment retail customers using K-Means clustering.
  • Customer segmentation helps businesses group their customers by purchasing behavior.
  • With clustering, companies can design personalized marketing plans and identify high-value clients.
  • You will analyze real retail transaction data to group customers by their spending patterns and frequency.
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
from sklearn.cluster import KMeans
from sklearn.preprocessing import StandardScaler
import warnings
warnings.filterwarnings('ignore')

Core Concepts: Retail Data & Sales Metrics#

  • Retail datasets usually include transactions, orders, products, and customer details.
  • Important sales metrics include revenue (Quantity x Price), order frequency, and recency.
  • Revenue is often miscalculated by forgetting to multiply price and quantity.
  • Grouping by the wrong column, such as by product instead of customer, is common among beginners.
  • Maintaining correct data types and formatting is key to accurate analytics.
# Load a real online retail transaction dataset
url = 'https://archive.ics.uci.edu/ml/machine-learning-databases/00502/online_retail_II.xlsx'
df = pd.read_excel(url, sheet_name='Year 2010-2011')
df['InvoiceDate'] = pd.to_datetime(df['InvoiceDate'])
print(df.shape)
print(df.head(3))
(541910, 8)
  Invoice StockCode                         Description  Quantity  \
0  536365    85123A  WHITE HANGING HEART T-LIGHT HOLDER         6   
1  536365     71053                 WHITE METAL LANTERN         6   
2  536365    84406B      CREAM CUPID HEARTS COAT HANGER         8   

          InvoiceDate  Price  Customer ID         Country  
0 2010-12-01 08:26:00   2.55      17850.0  United Kingdom  
1 2010-12-01 08:26:00   3.39      17850.0  United Kingdom  
2 2010-12-01 08:26:00   2.75      17850.0  United Kingdom  
# Check for missing values in Customer ID and Price/Quantity
missing_customers = df['Customer ID'].isna().sum()
missing_quantities = df['Quantity'].isna().sum()
missing_prices = df['Price'].isna().sum()
print('Missing Customer IDs:', missing_customers)
print('Missing Quantities:', missing_quantities)
print('Missing Prices:', missing_prices)
Missing Customer IDs: 135080
Missing Quantities: 0
Missing Prices: 0
# Remove transactions with missing Customer ID or negative Quantity/Price
df = df.dropna(subset=['Customer ID'])
df = df[(df['Quantity'] > 0) & (df['Price'] > 0)]
print('After cleaning:', df.shape)
After cleaning: (397885, 8)
# Calculate the total revenue per transaction
df['Revenue'] = df['Quantity'] * df['Price']
print(df[['Customer ID','Revenue']].head(3))
   Customer ID  Revenue
0      17850.0    15.30
1      17850.0    20.34
2      17850.0    22.00
# Example 1: Compute total revenue for each customer
customer_revenue = df.groupby('Customer ID')['Revenue'].sum().reset_index()
print(customer_revenue.head(3))
   Customer ID   Revenue
0      12346.0  77183.60
1      12347.0   4310.00
2      12348.0   1797.24
# Example 2: Compute number of purchases per customer
customer_freq = df.groupby('Customer ID')['Invoice'].nunique().reset_index(name='NumPurchases')
print(customer_freq.head(3))
   Customer ID  NumPurchases
0      12346.0             1
1      12347.0             7
2      12348.0             4
# Example 3: Compute recency for each customer (days since last purchase)
latest_date = df['InvoiceDate'].max()
recency = df.groupby('Customer ID')['InvoiceDate'].max().reset_index()
recency['RecencyDays'] = (latest_date - recency['InvoiceDate']).dt.days
recency = recency[['Customer ID','RecencyDays']]
print(recency.head(3))
   Customer ID  RecencyDays
0      12346.0          325
1      12347.0            1
2      12348.0           74
# Example 4 (intermediate): Merge revenue, frequency, and recency into one dataset
rfm = customer_revenue.merge(customer_freq, on='Customer ID')
rfm = rfm.merge(recency, on='Customer ID')
print(rfm.head(3))
   Customer ID   Revenue  NumPurchases  RecencyDays
0      12346.0  77183.60             1          325
1      12347.0   4310.00             7            1
2      12348.0   1797.24             4           74
# Example 5 (intermediate): Visualize distribution of RFM features
fig, axes = plt.subplots(1, 3, figsize=(15,4))
axes[0].hist(rfm['Revenue'], bins=30, color='skyblue')
axes[0].set_title('Customer Revenue Distribution')
axes[1].hist(rfm['NumPurchases'], bins=30, color='salmon')
axes[1].set_title('Purchase Frequency')
axes[2].hist(rfm['RecencyDays'], bins=30, color='lightgreen')
axes[2].set_title('Recency (Days)')
plt.tight_layout()
plt.show()
No description has been provided for this image
# Example 6 (intermediate): Scale RFM features for K-Means clustering
scaler = StandardScaler()
rfm_scaled = scaler.fit_transform(rfm[['Revenue','NumPurchases','RecencyDays']])
print('Means after scaling:', rfm_scaled.mean(axis=0))
print('Stds after scaling:', rfm_scaled.std(axis=0))
Means after scaling: [8.18975030e-18 1.80174507e-17 2.70261760e-17]
Stds after scaling: [1. 1. 1.]
# Example 7 (intermediate): Elbow method to choose number of clusters
wcss = []
for i in range(1, 11):
    kmeans = KMeans(n_clusters=i, random_state=42)
    kmeans.fit(rfm_scaled)
    wcss.append(kmeans.inertia_)
plt.plot(range(1,11), wcss, marker='o')
plt.xlabel('Number of clusters')
plt.ylabel('WCSS (Within-Cluster Sum of Squares)')
plt.title('Elbow Method for Finding Optimal K')
plt.show()
No description has been provided for this image
# Example 8 (advanced): Run K-Means with the chosen number of clusters (e.g., 4)
k = 4
kmeans = KMeans(n_clusters=k, random_state=42)
clusters = kmeans.fit_predict(rfm_scaled)
rfm['Cluster'] = clusters
print(rfm.head(3))
   Customer ID   Revenue  NumPurchases  RecencyDays  Cluster
0      12346.0  77183.60             1          325        3
1      12347.0   4310.00             7            1        0
2      12348.0   1797.24             4           74        0
# Example 9 (advanced): Profile customer clusters
cluster_summary = rfm.groupby('Cluster').agg({
    'Revenue': 'mean',
    'NumPurchases': 'mean',
    'RecencyDays': 'mean',
    'Customer ID': 'count'
}).rename(columns={'Customer ID': 'NumCustomers'}).reset_index()
print(cluster_summary)
   Cluster        Revenue  NumPurchases  RecencyDays  NumCustomers
0        0    1359.055178      3.682711    42.702685          3054
1        1     480.617480      1.552015   247.075914          1067
2        2  127338.313846     82.538462     6.384615            13
3        3   12709.090490     22.333333    14.500000           204
# Example 10 (advanced): Visualize clusters in 2D (using two principal features)
plt.figure(figsize=(8,6))
plt.scatter(rfm['Revenue'], rfm['RecencyDays'], c=rfm['Cluster'], cmap='tab10', alpha=0.7)
plt.xlabel('Revenue')
plt.ylabel('RecencyDays')
plt.title('Customer Clusters by Revenue and Recency')
plt.colorbar(label='Cluster')
plt.show()
No description has been provided for this image
# Error handling: What if a groupby fails due to missing columns?
try:
    df.groupby('NonExistentColumn').sum()
except Exception as e:
    print('Error:', e)
Error: 'NonExistentColumn'
# Error handling: Incorrect aggregation can break customer analysis
try:
    df['Total'] = df['Quantity'] + df['Price']  # Incorrect: adding, not multiplying
    print(df[['Quantity','Price','Total']].head(3))
except Exception as e:
    print('Error:', e)
   Quantity  Price  Total
0         6   2.55   8.55
1         6   3.39   9.39
2         8   2.75  10.75
# Error handling: Grouping by Product instead of Customer
product_agg = df.groupby('Description')['Revenue'].sum().reset_index()
print(product_agg.head(3))
                      Description  Revenue
0   4 PURPLE FLOCK DINNER CANDLES   270.76
1   50'S CHRISTMAS GIFT BAG LARGE  2272.25
2               DOLLY GIRL BEAKER  2759.50
# Best practices: Always validate cluster quality by profiling each cluster
for cluster in cluster_summary['Cluster']:
    print(f"Cluster {cluster}: Avg Revenue ${cluster_summary.loc[cluster,'Revenue']:.2f}, "
          f"Avg Purchases {cluster_summary.loc[cluster,'NumPurchases']:.1f}, "
          f"Recency {cluster_summary.loc[cluster,'RecencyDays']:.1f} days, "
          f"Customers: {cluster_summary.loc[cluster,'NumCustomers']}")
Cluster 0: Avg Revenue $1359.06, Avg Purchases 3.7, Recency 42.7 days, Customers: 3054
Cluster 1: Avg Revenue $480.62, Avg Purchases 1.6, Recency 247.1 days, Customers: 1067
Cluster 2: Avg Revenue $127338.31, Avg Purchases 82.5, Recency 6.4 days, Customers: 13
Cluster 3: Avg Revenue $12709.09, Avg Purchases 22.3, Recency 14.5 days, Customers: 204

Best Practices: Customer Segmentation Patterns#

  • Use clear and interpretable features: revenue, frequency, and recency are a solid foundation.
  • Standardize all feature values before clustering (especially when units differ).
  • Visualize both distributions and cluster assignments for actionable insights.
  • Profile and describe each cluster to discover what makes each group unique.
# End-to-end retail analytics problem: Identify and describe "VIP" customers
# Definition: VIPs = customers in the cluster with highest average revenue
vip_cluster = cluster_summary.loc[cluster_summary['Revenue'].idxmax(), 'Cluster']
vip_customers = rfm[rfm['Cluster'] == vip_cluster]
print(f"VIP customer count: {len(vip_customers)}")
print(vip_customers[['Customer ID', 'Revenue', 'NumPurchases', 'RecencyDays']].head())
VIP customer count: 13
      Customer ID    Revenue  NumPurchases  RecencyDays
55        12415.0  124914.53            21           23
326       12748.0   33719.73           209            0
562       13089.0   58825.83            97            2
1333      14156.0  117379.63            55            9
1689      14646.0  280206.02            73            1

Congratulations! Keep Practicing Advanced Retail Analytics#

  • Try different feature combinations or clustering algorithms to segment your customers.
  • For further learning, search YouTube for "Retail Customer Segmentation Python".
  • Share your findings and cluster plots with your study group or supervisor.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.