Lesson 30 · Market Research Analytics in Python
Customer Clustering Using K-Means
In this lesson, we will learn how to segment customers into meaningful groups using K-Means clustering. Customer segmentation helps businesses target…
- CourseMarket Research Analytics in Python
- Lesson30 of 56
- Video21 min
- FormatJupyter notebook · 19 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbCustomer Clustering Using K-Means#
- In this lesson, we will learn how to segment customers into meaningful groups using K-Means clustering.
- Customer segmentation helps businesses target marketing, improve products, and personalize communication.
- You will analyze real-world customer datasets, group customers by behavior, and discover business insights.
- The goal is to produce clear customer segments that enable better decision-making.
import pandas as pd
import numpy as np
import openml
from sklearn.preprocessing import StandardScaler, LabelEncoder
from sklearn.cluster import KMeans
import matplotlib.pyplot as plt
import seaborn as sns
import warnings
warnings.filterwarnings('ignore')
Understanding Customer Data for Clustering#
- Customer datasets commonly include demographics, purchase behavior, and feedback.
- Clustering works best with structured, numeric, or carefully prepared categorical data.
- Typical columns: age, gender, spending, region, tenure, or satisfaction.
- Beginners often use raw data without scaling or encoding, which hurts clustering accuracy.
- Segmentation mistakes include ignoring categorical data, failing to handle missing values, and misreading variable types.
url = 'https://archive.ics.uci.edu/ml/machine-learning-databases/00502/online_retail_II.xlsx'
df = pd.read_excel(url, sheet_name='Year 2010-2011')
df['InvoiceDate'] = pd.to_datetime(df['InvoiceDate'])
print(df.shape)
print(df.head(3))
customer_df = df.groupby('Customer ID').agg({
'Invoice': 'nunique',
'Quantity': 'sum',
'Price': 'sum'
}).reset_index().rename(columns={'Invoice':'NumPurchases','Quantity':'TotalQty','Price':'TotalSpent'})
print(customer_df.head(3))
print('Missing values per column:')
print(customer_df.isnull().sum())
customer_df = customer_df.dropna()
print('Customer sample after dropping missing:')
print(customer_df.head(3))
scaler = StandardScaler()
features = ['NumPurchases', 'TotalQty', 'TotalSpent']
X_scaled = scaler.fit_transform(customer_df[features])
print('Means after scaling:', X_scaled.mean(axis=0))
inertia = []
K_range = range(1, 10)
for k in K_range:
km = KMeans(n_clusters=k, random_state=42)
km.fit(X_scaled)
inertia.append(km.inertia_)
plt.figure(figsize=(7, 4))
plt.plot(K_range, inertia, marker='o')
plt.xlabel('Number of clusters (k)')
plt.ylabel('Inertia')
plt.title('Elbow Method for Optimal k')
plt.show()
best_k = 3
kmeans = KMeans(n_clusters=best_k, random_state=42)
clusters = kmeans.fit_predict(X_scaled)
customer_df['Cluster'] = clusters
print(customer_df['Cluster'].value_counts())
sns.pairplot(customer_df, hue='Cluster', vars=features, palette='Set1')
plt.suptitle('Customer Clusters by Behavior', y=1.02)
plt.show()
summary = customer_df.groupby('Cluster')[features].mean()
print('Average behavior per cluster:')
print(summary)
summary.plot(kind='bar', figsize=(8,5))
plt.title('Average Customer Metrics per Segment')
plt.ylabel('Standardized Value')
plt.xlabel('Cluster Label')
plt.legend(loc='best')
plt.show()
dataset = openml.datasets.get_dataset(42178)
df_survey, _, _, _ = dataset.get_data(dataset_format='dataframe')
print(df_survey.shape)
print(df_survey.head(3))
cat_columns = ['gender', 'Partner', 'Dependents', 'InternetService', 'Contract', 'PaymentMethod']
survey_enc = df_survey.copy()
for col in cat_columns:
survey_enc[col] = LabelEncoder().fit_transform(survey_enc[col].astype(str))
print(survey_enc[cat_columns].head(3))
survey_features = ['gender', 'SeniorCitizen', 'Partner', 'tenure', 'MonthlyCharges']
survey_scaled = StandardScaler().fit_transform(survey_enc[survey_features])
survey_kmeans = KMeans(n_clusters=3, random_state=42)
survey_clusters = survey_kmeans.fit_predict(survey_scaled)
df_survey['Segment'] = survey_clusters
print(df_survey.groupby('Segment').size())
# Customer NPS Clustering Example
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
nps_df = pd.DataFrame({
'CustomerID': range(1, 501),
'Age': np.random.randint(18, 70, 500),
'Region': np.random.choice(['North', 'South', 'East', 'West'], 500),
'NPS_Score': np.random.randint(0, 11, 500)
})
X_nps = nps_df[['Age', 'NPS_Score']]
scaler = StandardScaler()
X_nps_scaled = scaler.fit_transform(X_nps)
kmeans_nps = KMeans(n_clusters=4, random_state=42)
nps_clusters = kmeans_nps.fit_predict(X_nps_scaled)
nps_df['Cluster'] = nps_clusters
cluster_summary = nps_df.groupby('Cluster')[['NPS_Score', 'Age']].mean()
print(cluster_summary)
plt.figure(figsize=(8,5))
plt.scatter(nps_df['Age'], nps_df['NPS_Score'], c=nps_df['Cluster'], cmap='viridis', alpha=0.6)
plt.title('Customer Segments Based on Age and NPS Score')
plt.xlabel('Age')
plt.ylabel('NPS Score')
plt.colorbar(label='Cluster')
plt.show()
missing_before = df_survey.isnull().sum().sum()
df_survey = df_survey.fillna(method='ffill').fillna(method='bfill')
missing_after = df_survey.isnull().sum().sum()
print('Missing before:', missing_before)
print('Missing after:', missing_after)
try:
# Intentionally grouping without resetting index
err_df = df.groupby('Country').agg({'Quantity':'sum'})
print('Grouped:', err_df.head(2))
except Exception as e:
print('Grouping error:', e)
nps_values = np.random.randint(0, 11, 20)
labels = ['Detractor' if s < 7 else 'Promoter' if s > 8 else 'Passive' for s in nps_values]
print('NPS values:', nps_values)
print('Labeled interpretation:', labels)
Market Research Analytics Patterns and Best Practices#
- Always preprocess by scaling and encoding before clustering.
- Use segmentation to identify marketing or operational opportunities.
- Cross-tabulate segments with customer outcomes (e.g., churn or satisfaction).
- Build index scores for composite measures of loyalty or satisfaction.
- Analyze trends over time to see if segments behave differently.
filtered = customer_df[customer_df['TotalSpent'] > 0]
top_segment = filtered.groupby('Cluster')['TotalSpent'].mean().idxmax()
print(f'The most valuable segment is Cluster {top_segment}:')
segment_members = filtered[filtered['Cluster'] == top_segment]
print('Sample high-value customers:')
print(segment_members[['Customer ID','TotalSpent']].head())
report = 'Customer Segmentation using K-Means\n\n'
report += f'Total customers: {len(filtered)}\n'
for cid, rec in segment_members.head(3).iterrows():
report += f'CustomerID: {rec[0]}, TotalSpent: {rec[4]:.2f}\n'
with open('cluster_report.txt', 'w') as f:
f.write(report)
print('Report saved.')
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



