Lesson 57 · Market Research Analytics in Python
Customer Segmentation Case Study Using Python for Market Research Analytics
In this lesson, we solve a real-world customer segmentation problem for market research. Segmenting customers allows businesses to tailor marketing and…
- CourseMarket Research Analytics in Python
- Lesson57 of 56
- Video21 min
- FormatJupyter notebook · 25 code cells
What you'll learn
- Conceptual background for market research data
- Beginner Example: Loading a real customer survey dataset
- Beginner Example: Summarizing customer demographics
- Beginner Example: Checking for missing data
- Intermediate Example: Exploring service usage patterns
- Intermediate Example: Segmenting by monthly charges
- Intermediate Example: Relationship between contract type and churn
- Advanced Example: Standardizing data for clustering
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbCustomer Segmentation Case Study#
- In this lesson, we solve a real-world customer segmentation problem for market research.
- Segmenting customers allows businesses to tailor marketing and services to different customer groups.
- You will learn to load, explore, segment, and analyze customer data to uncover actionable insights.
- By the end, you will be able to perform demographic segmentation and find key customer groups.
import warnings
warnings.filterwarnings('ignore')
import pandas as pd
import numpy as np
import openml
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
import matplotlib.pyplot as plt
import seaborn as sns
Conceptual background for market research data#
- Market research data includes survey responses, customer demographics, purchase records, and transactional data.
- Each row typically represents a customer, with columns for attributes like age, region, behavior, and satisfaction.
- Beginners often mistake categorical values for numeric, overlook missing values, or aggregate data incorrectly.
- Interpreting scales, like Likert or NPS, incorrectly can result in misleading insights.
Beginner Example: Loading a real customer survey dataset#
- We will use OpenML's customer satisfaction dataset to practice segmentation.
- Columns include customer demographics, services used, and churn status.
dataset = openml.datasets.get_dataset(42178)
df, _, _, _ = dataset.get_data(dataset_format='dataframe')
print(df.shape)
print(df.head(3))
Beginner Example: Summarizing customer demographics#
- Understanding the distribution of customer demographics is the first step in segmentation.
- Let's view gender and age-related columns.
print(df['gender'].value_counts())
print(df['SeniorCitizen'].value_counts())
print(df[['gender','SeniorCitizen','Partner','Dependents']].describe(include='all'))
Beginner Example: Checking for missing data#
- Missing values are common in surveys and must be handled to avoid biased segments.
- We will check if there are any missing responses in the dataset.
print(df.isnull().sum())
Intermediate Example: Exploring service usage patterns#
- Segmentation can be based on the services customers use.
- Let's compare use of internet and phone services.
service_counts = df.groupby(['InternetService', 'PhoneService']).size().unstack()
print(service_counts)
ax = service_counts.plot(kind='bar', stacked=True, figsize=(8,5))
plt.title('Customer Segments by Internet and Phone Service')
plt.ylabel('Number of Customers')
plt.xticks(rotation=0)
plt.tight_layout()
plt.show()
Intermediate Example: Segmenting by monthly charges#
- Customers can be grouped into segments based on spending.
- Let's create spending groups using MonthlyCharges.
df['SpendingSegment'] = pd.cut(df['MonthlyCharges'], bins=[0,40,80, df['MonthlyCharges'].max()], labels=['Low','Medium','High'])
print(df['SpendingSegment'].value_counts())
sns.countplot(x='SpendingSegment', data=df, palette='Set2')
plt.title('Customer Segments by Monthly Charges')
plt.xlabel('Spending Segment')
plt.ylabel('Count')
plt.show()
Intermediate Example: Relationship between contract type and churn#
- Business insight: Churn can be higher for certain contract types.
- Let's analyze churn rate across contract segments.
churn_by_contract = df.groupby('Contract')['Churn'].value_counts(normalize=True).unstack().fillna(0)
print(churn_by_contract)
churn_by_contract.plot(kind='bar', stacked=True, color=['green','orange'], figsize=(8,5))
plt.title('Churn Rate Across Contract Types')
plt.ylabel('Proportion of Customers')
plt.xticks(rotation=0)
plt.tight_layout()
plt.show()
Advanced Example: Standardizing data for clustering#
- K-means segmentation requires standardized data for reliable results.
- We will select numeric columns and standardize them.
features = ['MonthlyCharges','TotalCharges','tenure']
df_numeric = df[features].replace(' ', np.nan).fillna(0).astype(float)
scaler = StandardScaler()
df_scaled = scaler.fit_transform(df_numeric)
Advanced Example: Finding optimal number of customer segments#
- K-means requires us to pick the right number of segments (k).
- We use the elbow method to visually determine the optimal k.
inertia = []
np.random.seed(42)
for k in range(1,9):
model = KMeans(n_clusters=k, random_state=42)
model.fit(df_scaled)
inertia.append(model.inertia_)
plt.plot(range(1,9), inertia, marker='o')
plt.title('Elbow Method For Optimal k')
plt.xlabel('Number of Clusters')
plt.ylabel('Inertia')
plt.show()
kmeans = KMeans(n_clusters=3, random_state=42)
df['Cluster'] = kmeans.fit_predict(df_scaled)
print(df['Cluster'].value_counts())
Advanced Example: Profiling customer clusters#
- Segment profiles tell us how clusters differ in business variables.
- Let's summarize average values of each segment.
# Safe Cluster Profiling with Numeric Enforcement
df[features] = df[features].apply(pd.to_numeric, errors='coerce')
df = df.dropna(subset=['Cluster'])
cluster_profile = (
df
.groupby('Cluster')[features]
.mean()
)
print(cluster_profile)
sns.boxplot(x='Cluster', y='MonthlyCharges', data=df)
plt.title('Monthly Charges Distribution by Cluster')
plt.xlabel('Customer Segment')
plt.ylabel('Monthly Charges')
plt.show()
Error handling: Dealing with missing survey responses#
- Missing values can break segmentation algorithms if not addressed.
- Let's simulate missing data in MonthlyCharges and see the effect.
df_missing = df.copy()
df_missing.loc[df_missing.sample(frac=0.05, random_state=42).index, 'MonthlyCharges'] = np.nan
print(df_missing['MonthlyCharges'].isnull().sum())
try:
df_missing['MonthlyCharges'].astype(float).mean()
except Exception as e:
print('Error:', e)
Error handling: Grouping by the wrong variable#
- Grouping by the wrong fields causes misleading business insights.
- Let's see what happens if you aggregate by a non-segmentation variable.
wrong_group = df.groupby('PaymentMethod')['MonthlyCharges'].mean()
print(wrong_group)
Error handling: Misinterpreting Likert or NPS survey scores#
- Scores like 0-10 NPS measure satisfaction but should not be averaged blindly.
- Let us see an appropriate and inappropriate way to analyze NPS-like data.
nps_df = pd.DataFrame({'Score':[0, 4, 6, 8, 9, 10]})
mean_score = nps_df['Score'].mean()
print(f'Average score (not the best for NPS): {mean_score}')
promoters = (nps_df['Score'] >= 9).sum()
detractors = (nps_df['Score'] <= 6).sum()
nps = (promoters - detractors) / len(nps_df) * 100
print(f'Calculated NPS score (correct): {nps}')
Best Practices: Segmentation, Cross-tabulation, and Trend Analysis#
- Use segmentation to define actionable customer groups.
- Employ cross-tabs for comparing segment behaviors.
- Use trend analysis over time for strategic decisions.
- Always check for missing data and standardize features for clustering.
cross_tab = pd.crosstab(df['SpendingSegment'], df['Churn'])
print(cross_tab)
segment_trend = df.groupby(['SpendingSegment','Contract']).size().unstack()
segment_trend.plot(kind='bar', stacked=True)
plt.title('Contract Types by Spending Segment')
plt.xlabel('Spending Segment')
plt.ylabel('Number of Customers')
plt.tight_layout()
plt.show()
End-to-end Market Research Problem: Segmenting and targeting customers#
- Let's walk through the process from data to actionable insight.
- We will identify key segments and recommend a strategy.
segment_counts = df.groupby(['Cluster','SpendingSegment']).size().unstack(fill_value=0)
print(segment_counts)
high_value = df[(df['SpendingSegment']=='High') & (df['Churn']=='No')]
print('Number of loyal high-value customers:', high_value.shape[0])
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



