Mathew K Analytics

Lesson 30 · Market Research Analytics in Python

Customer Clustering Using K-Means

In this lesson, we will learn how to segment customers into meaningful groups using K-Means clustering. Customer segmentation helps businesses target…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Customer Clustering Using K-Means#

  • In this lesson, we will learn how to segment customers into meaningful groups using K-Means clustering.
  • Customer segmentation helps businesses target marketing, improve products, and personalize communication.
  • You will analyze real-world customer datasets, group customers by behavior, and discover business insights.
  • The goal is to produce clear customer segments that enable better decision-making.
import pandas as pd
import numpy as np
import openml
from sklearn.preprocessing import StandardScaler, LabelEncoder
from sklearn.cluster import KMeans
import matplotlib.pyplot as plt
import seaborn as sns
import warnings
warnings.filterwarnings('ignore')

Understanding Customer Data for Clustering#

  • Customer datasets commonly include demographics, purchase behavior, and feedback.
  • Clustering works best with structured, numeric, or carefully prepared categorical data.
  • Typical columns: age, gender, spending, region, tenure, or satisfaction.
  • Beginners often use raw data without scaling or encoding, which hurts clustering accuracy.
  • Segmentation mistakes include ignoring categorical data, failing to handle missing values, and misreading variable types.
url = 'https://archive.ics.uci.edu/ml/machine-learning-databases/00502/online_retail_II.xlsx'
df = pd.read_excel(url, sheet_name='Year 2010-2011')
df['InvoiceDate'] = pd.to_datetime(df['InvoiceDate'])
print(df.shape)
print(df.head(3))
(541910, 8)
  Invoice StockCode                         Description  Quantity  \
0  536365    85123A  WHITE HANGING HEART T-LIGHT HOLDER         6   
1  536365     71053                 WHITE METAL LANTERN         6   
2  536365    84406B      CREAM CUPID HEARTS COAT HANGER         8   

          InvoiceDate  Price  Customer ID         Country  
0 2010-12-01 08:26:00   2.55      17850.0  United Kingdom  
1 2010-12-01 08:26:00   3.39      17850.0  United Kingdom  
2 2010-12-01 08:26:00   2.75      17850.0  United Kingdom  
customer_df = df.groupby('Customer ID').agg({
    'Invoice': 'nunique',
    'Quantity': 'sum',
    'Price': 'sum'
}).reset_index().rename(columns={'Invoice':'NumPurchases','Quantity':'TotalQty','Price':'TotalSpent'})
print(customer_df.head(3))
   Customer ID  NumPurchases  TotalQty  TotalSpent
0      12346.0             2         0        2.08
1      12347.0             7      2458      481.21
2      12348.0             4      2341      178.71
print('Missing values per column:')
print(customer_df.isnull().sum())
customer_df = customer_df.dropna()
print('Customer sample after dropping missing:')
print(customer_df.head(3))
Missing values per column:
Customer ID     0
NumPurchases    0
TotalQty        0
TotalSpent      0
dtype: int64
Customer sample after dropping missing:
   Customer ID  NumPurchases  TotalQty  TotalSpent
0      12346.0             2         0        2.08
1      12347.0             7      2458      481.21
2      12348.0             4      2341      178.71
scaler = StandardScaler()
features = ['NumPurchases', 'TotalQty', 'TotalSpent']
X_scaled = scaler.fit_transform(customer_df[features])
print('Means after scaling:', X_scaled.mean(axis=0))
Means after scaling: [ 1.95025454e-17 -1.78773332e-17 -1.95025454e-17]
inertia = []
K_range = range(1, 10)
for k in K_range:
    km = KMeans(n_clusters=k, random_state=42)
    km.fit(X_scaled)
    inertia.append(km.inertia_)
plt.figure(figsize=(7, 4))
plt.plot(K_range, inertia, marker='o')
plt.xlabel('Number of clusters (k)')
plt.ylabel('Inertia')
plt.title('Elbow Method for Optimal k')
plt.show()
No description has been provided for this image
best_k = 3
kmeans = KMeans(n_clusters=best_k, random_state=42)
clusters = kmeans.fit_predict(X_scaled)
customer_df['Cluster'] = clusters
print(customer_df['Cluster'].value_counts())
Cluster
0    4054
2     300
1      18
Name: count, dtype: int64
sns.pairplot(customer_df, hue='Cluster', vars=features, palette='Set1')
plt.suptitle('Customer Clusters by Behavior', y=1.02)
plt.show()
No description has been provided for this image
summary = customer_df.groupby('Cluster')[features].mean()
print('Average behavior per cluster:')
print(summary)
summary.plot(kind='bar', figsize=(8,5))
plt.title('Average Customer Metrics per Segment')
plt.ylabel('Standardized Value')
plt.xlabel('Cluster Label')
plt.legend(loc='best')
plt.show()
Average behavior per cluster:
         NumPurchases      TotalQty    TotalSpent
Cluster                                          
0            3.409225    606.669956    198.195095
1           87.055556  49971.666667  13656.025000
2           22.673333   5159.863333   1195.155333
No description has been provided for this image
dataset = openml.datasets.get_dataset(42178)
df_survey, _, _, _ = dataset.get_data(dataset_format='dataframe')
print(df_survey.shape)
print(df_survey.head(3))
(7043, 20)
   gender  SeniorCitizen Partner Dependents  tenure PhoneService  \
0  Female              0     Yes         No       1           No   
1    Male              0      No         No      34          Yes   
2    Male              0      No         No       2          Yes   

      MultipleLines InternetService OnlineSecurity OnlineBackup  \
0  No phone service             DSL             No          Yes   
1                No             DSL            Yes           No   
2                No             DSL            Yes          Yes   

  DeviceProtection TechSupport StreamingTV StreamingMovies        Contract  \
0               No          No          No              No  Month-to-month   
1              Yes          No          No              No        One year   
2               No          No          No              No  Month-to-month   

  PaperlessBilling     PaymentMethod  MonthlyCharges TotalCharges Churn  
0              Yes  Electronic check           29.85        29.85    No  
1               No      Mailed check           56.95       1889.5    No  
2              Yes      Mailed check           53.85       108.15   Yes  
cat_columns = ['gender', 'Partner', 'Dependents', 'InternetService', 'Contract', 'PaymentMethod']
survey_enc = df_survey.copy()
for col in cat_columns:
    survey_enc[col] = LabelEncoder().fit_transform(survey_enc[col].astype(str))
print(survey_enc[cat_columns].head(3))
   gender  Partner  Dependents  InternetService  Contract  PaymentMethod
0       0        1           0                0         0              2
1       1        0           0                0         1              3
2       1        0           0                0         0              3
survey_features = ['gender', 'SeniorCitizen', 'Partner', 'tenure', 'MonthlyCharges']
survey_scaled = StandardScaler().fit_transform(survey_enc[survey_features])
survey_kmeans = KMeans(n_clusters=3, random_state=42)
survey_clusters = survey_kmeans.fit_predict(survey_scaled)
df_survey['Segment'] = survey_clusters
print(df_survey.groupby('Segment').size())
Segment
0    2065
1    2903
2    2075
dtype: int64
# Customer NPS Clustering Example
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
nps_df = pd.DataFrame({
    'CustomerID': range(1, 501),
    'Age': np.random.randint(18, 70, 500),
    'Region': np.random.choice(['North', 'South', 'East', 'West'], 500),
    'NPS_Score': np.random.randint(0, 11, 500)
})
X_nps = nps_df[['Age', 'NPS_Score']]
scaler = StandardScaler()
X_nps_scaled = scaler.fit_transform(X_nps)
kmeans_nps = KMeans(n_clusters=4, random_state=42)
nps_clusters = kmeans_nps.fit_predict(X_nps_scaled)
nps_df['Cluster'] = nps_clusters
cluster_summary = nps_df.groupby('Cluster')[['NPS_Score', 'Age']].mean()
print(cluster_summary)
plt.figure(figsize=(8,5))
plt.scatter(nps_df['Age'], nps_df['NPS_Score'], c=nps_df['Cluster'], cmap='viridis', alpha=0.6)
plt.title('Customer Segments Based on Age and NPS Score')
plt.xlabel('Age')
plt.ylabel('NPS Score')
plt.colorbar(label='Cluster')
plt.show()
         NPS_Score        Age
Cluster                      
0         1.750000  55.046296
1         7.376712  57.109589
2         2.627907  30.364341
3         8.000000  31.128205
No description has been provided for this image
missing_before = df_survey.isnull().sum().sum()
df_survey = df_survey.fillna(method='ffill').fillna(method='bfill')
missing_after = df_survey.isnull().sum().sum()
print('Missing before:', missing_before)
print('Missing after:', missing_after)
Missing before: 0
Missing after: 0
try:
    # Intentionally grouping without resetting index
    err_df = df.groupby('Country').agg({'Quantity':'sum'})
    print('Grouped:', err_df.head(2))
except Exception as e:
    print('Grouping error:', e)
Grouped:            Quantity
Country            
Australia     83653
Austria        4827
nps_values = np.random.randint(0, 11, 20)
labels = ['Detractor' if s < 7 else 'Promoter' if s > 8 else 'Passive' for s in nps_values]
print('NPS values:', nps_values)
print('Labeled interpretation:', labels)
NPS values: [ 4  7  4  8  7  0  6  9  5  4  9  9  9  4  7  9  9  1 10  4]
Labeled interpretation: ['Detractor', 'Passive', 'Detractor', 'Passive', 'Passive', 'Detractor', 'Detractor', 'Promoter', 'Detractor', 'Detractor', 'Promoter', 'Promoter', 'Promoter', 'Detractor', 'Passive', 'Promoter', 'Promoter', 'Detractor', 'Promoter', 'Detractor']

Market Research Analytics Patterns and Best Practices#

  • Always preprocess by scaling and encoding before clustering.
  • Use segmentation to identify marketing or operational opportunities.
  • Cross-tabulate segments with customer outcomes (e.g., churn or satisfaction).
  • Build index scores for composite measures of loyalty or satisfaction.
  • Analyze trends over time to see if segments behave differently.
filtered = customer_df[customer_df['TotalSpent'] > 0]
top_segment = filtered.groupby('Cluster')['TotalSpent'].mean().idxmax()
print(f'The most valuable segment is Cluster {top_segment}:')
segment_members = filtered[filtered['Cluster'] == top_segment]
print('Sample high-value customers:')
print(segment_members[['Customer ID','TotalSpent']].head())
The most valuable segment is Cluster 1:
Sample high-value customers:
      Customer ID  TotalSpent
55        12415.0     2499.82
328       12744.0    25108.89
330       12748.0    15115.60
568       13089.0     5166.45
1005      13694.0     1163.81
report = 'Customer Segmentation using K-Means\n\n'
report += f'Total customers: {len(filtered)}\n'
for cid, rec in segment_members.head(3).iterrows():
    report += f'CustomerID: {rec[0]}, TotalSpent: {rec[4]:.2f}\n'
with open('cluster_report.txt', 'w') as f:
    f.write(report)
print('Report saved.')
Report saved.
 

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.