Mathew K Analytics

Lesson 57 · Market Research Analytics in Python

Customer Segmentation Case Study Using Python for Market Research Analytics

In this lesson, we solve a real-world customer segmentation problem for market research. Segmenting customers allows businesses to tailor marketing and…

⬇ Download notebookOpen in Colab ↗

📓 Full notebook

Download .ipynb

Customer Segmentation Case Study#

  • In this lesson, we solve a real-world customer segmentation problem for market research.
  • Segmenting customers allows businesses to tailor marketing and services to different customer groups.
  • You will learn to load, explore, segment, and analyze customer data to uncover actionable insights.
  • By the end, you will be able to perform demographic segmentation and find key customer groups.
import warnings
warnings.filterwarnings('ignore')
import pandas as pd
import numpy as np
import openml
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
import matplotlib.pyplot as plt
import seaborn as sns

Conceptual background for market research data#

  • Market research data includes survey responses, customer demographics, purchase records, and transactional data.
  • Each row typically represents a customer, with columns for attributes like age, region, behavior, and satisfaction.
  • Beginners often mistake categorical values for numeric, overlook missing values, or aggregate data incorrectly.
  • Interpreting scales, like Likert or NPS, incorrectly can result in misleading insights.

Beginner Example: Loading a real customer survey dataset#

  • We will use OpenML's customer satisfaction dataset to practice segmentation.
  • Columns include customer demographics, services used, and churn status.
dataset = openml.datasets.get_dataset(42178)
df, _, _, _ = dataset.get_data(dataset_format='dataframe')
print(df.shape)
print(df.head(3))
(7043, 20)
   gender  SeniorCitizen Partner Dependents  tenure PhoneService  \
0  Female              0     Yes         No       1           No   
1    Male              0      No         No      34          Yes   
2    Male              0      No         No       2          Yes   

      MultipleLines InternetService OnlineSecurity OnlineBackup  \
0  No phone service             DSL             No          Yes   
1                No             DSL            Yes           No   
2                No             DSL            Yes          Yes   

  DeviceProtection TechSupport StreamingTV StreamingMovies        Contract  \
0               No          No          No              No  Month-to-month   
1              Yes          No          No              No        One year   
2               No          No          No              No  Month-to-month   

  PaperlessBilling     PaymentMethod  MonthlyCharges TotalCharges Churn  
0              Yes  Electronic check           29.85        29.85    No  
1               No      Mailed check           56.95       1889.5    No  
2              Yes      Mailed check           53.85       108.15   Yes  

Beginner Example: Summarizing customer demographics#

  • Understanding the distribution of customer demographics is the first step in segmentation.
  • Let's view gender and age-related columns.
print(df['gender'].value_counts())
print(df['SeniorCitizen'].value_counts())
gender
Male      3555
Female    3488
Name: count, dtype: int64
SeniorCitizen
0    5901
1    1142
Name: count, dtype: int64
print(df[['gender','SeniorCitizen','Partner','Dependents']].describe(include='all'))
       gender  SeniorCitizen Partner Dependents
count    7043    7043.000000    7043       7043
unique      2            NaN       2          2
top      Male            NaN      No         No
freq     3555            NaN    3641       4933
mean      NaN       0.162147     NaN        NaN
std       NaN       0.368612     NaN        NaN
min       NaN       0.000000     NaN        NaN
25%       NaN       0.000000     NaN        NaN
50%       NaN       0.000000     NaN        NaN
75%       NaN       0.000000     NaN        NaN
max       NaN       1.000000     NaN        NaN

Beginner Example: Checking for missing data#

  • Missing values are common in surveys and must be handled to avoid biased segments.
  • We will check if there are any missing responses in the dataset.
print(df.isnull().sum())
gender              0
SeniorCitizen       0
Partner             0
Dependents          0
tenure              0
PhoneService        0
MultipleLines       0
InternetService     0
OnlineSecurity      0
OnlineBackup        0
DeviceProtection    0
TechSupport         0
StreamingTV         0
StreamingMovies     0
Contract            0
PaperlessBilling    0
PaymentMethod       0
MonthlyCharges      0
TotalCharges        0
Churn               0
dtype: int64

Intermediate Example: Exploring service usage patterns#

  • Segmentation can be based on the services customers use.
  • Let's compare use of internet and phone services.
service_counts = df.groupby(['InternetService', 'PhoneService']).size().unstack()
print(service_counts)
PhoneService        No     Yes
InternetService               
DSL              682.0  1739.0
Fiber optic        NaN  3096.0
No                 NaN  1526.0
ax = service_counts.plot(kind='bar', stacked=True, figsize=(8,5))
plt.title('Customer Segments by Internet and Phone Service')
plt.ylabel('Number of Customers')
plt.xticks(rotation=0)
plt.tight_layout()
plt.show()
No description has been provided for this image

Intermediate Example: Segmenting by monthly charges#

  • Customers can be grouped into segments based on spending.
  • Let's create spending groups using MonthlyCharges.
df['SpendingSegment'] = pd.cut(df['MonthlyCharges'], bins=[0,40,80, df['MonthlyCharges'].max()], labels=['Low','Medium','High'])
print(df['SpendingSegment'].value_counts())
SpendingSegment
High      2666
Medium    2539
Low       1838
Name: count, dtype: int64
sns.countplot(x='SpendingSegment', data=df, palette='Set2')
plt.title('Customer Segments by Monthly Charges')
plt.xlabel('Spending Segment')
plt.ylabel('Count')
plt.show()
No description has been provided for this image

Intermediate Example: Relationship between contract type and churn#

  • Business insight: Churn can be higher for certain contract types.
  • Let's analyze churn rate across contract segments.
churn_by_contract = df.groupby('Contract')['Churn'].value_counts(normalize=True).unstack().fillna(0)
print(churn_by_contract)
Churn                 No       Yes
Contract                          
Month-to-month  0.572903  0.427097
One year        0.887305  0.112695
Two year        0.971681  0.028319
churn_by_contract.plot(kind='bar', stacked=True, color=['green','orange'], figsize=(8,5))
plt.title('Churn Rate Across Contract Types')
plt.ylabel('Proportion of Customers')
plt.xticks(rotation=0)
plt.tight_layout()
plt.show()
No description has been provided for this image

Advanced Example: Standardizing data for clustering#

  • K-means segmentation requires standardized data for reliable results.
  • We will select numeric columns and standardize them.
features = ['MonthlyCharges','TotalCharges','tenure']
df_numeric = df[features].replace(' ', np.nan).fillna(0).astype(float)
scaler = StandardScaler()
df_scaled = scaler.fit_transform(df_numeric)

Advanced Example: Finding optimal number of customer segments#

  • K-means requires us to pick the right number of segments (k).
  • We use the elbow method to visually determine the optimal k.
inertia = []
np.random.seed(42)
for k in range(1,9):
    model = KMeans(n_clusters=k, random_state=42)
    model.fit(df_scaled)
    inertia.append(model.inertia_)
plt.plot(range(1,9), inertia, marker='o')
plt.title('Elbow Method For Optimal k')
plt.xlabel('Number of Clusters')
plt.ylabel('Inertia')
plt.show()
No description has been provided for this image
kmeans = KMeans(n_clusters=3, random_state=42)
df['Cluster'] = kmeans.fit_predict(df_scaled)
print(df['Cluster'].value_counts())
Cluster
0    2681
1    2201
2    2161
Name: count, dtype: int64

Advanced Example: Profiling customer clusters#

  • Segment profiles tell us how clusters differ in business variables.
  • Let's summarize average values of each segment.
# Safe Cluster Profiling with Numeric Enforcement
df[features] = df[features].apply(pd.to_numeric, errors='coerce')
df = df.dropna(subset=['Cluster'])
cluster_profile = (
    df
    .groupby('Cluster')[features]
    .mean()
)
print(cluster_profile)
         MonthlyCharges  TotalCharges     tenure
Cluster                                         
0             75.039276   1034.223038  13.256994
1             89.680032   5245.050704  58.559291
2             26.631444    809.388051  29.411846
sns.boxplot(x='Cluster', y='MonthlyCharges', data=df)
plt.title('Monthly Charges Distribution by Cluster')
plt.xlabel('Customer Segment')
plt.ylabel('Monthly Charges')
plt.show()
No description has been provided for this image

Error handling: Dealing with missing survey responses#

  • Missing values can break segmentation algorithms if not addressed.
  • Let's simulate missing data in MonthlyCharges and see the effect.
df_missing = df.copy()
df_missing.loc[df_missing.sample(frac=0.05, random_state=42).index, 'MonthlyCharges'] = np.nan
print(df_missing['MonthlyCharges'].isnull().sum())
352
try:
    df_missing['MonthlyCharges'].astype(float).mean()
except Exception as e:
    print('Error:', e)

Error handling: Grouping by the wrong variable#

  • Grouping by the wrong fields causes misleading business insights.
  • Let's see what happens if you aggregate by a non-segmentation variable.
wrong_group = df.groupby('PaymentMethod')['MonthlyCharges'].mean()
print(wrong_group)
PaymentMethod
Bank transfer (automatic)    67.192649
Credit card (automatic)      66.512385
Electronic check             76.255814
Mailed check                 43.917060
Name: MonthlyCharges, dtype: float64

Error handling: Misinterpreting Likert or NPS survey scores#

  • Scores like 0-10 NPS measure satisfaction but should not be averaged blindly.
  • Let us see an appropriate and inappropriate way to analyze NPS-like data.
nps_df = pd.DataFrame({'Score':[0, 4, 6, 8, 9, 10]})
mean_score = nps_df['Score'].mean()
print(f'Average score (not the best for NPS): {mean_score}')
promoters = (nps_df['Score'] >= 9).sum()
detractors = (nps_df['Score'] <= 6).sum()
nps = (promoters - detractors) / len(nps_df) * 100
print(f'Calculated NPS score (correct): {nps}')
Average score (not the best for NPS): 6.166666666666667
Calculated NPS score (correct): -16.666666666666664

Best Practices: Segmentation, Cross-tabulation, and Trend Analysis#

  • Use segmentation to define actionable customer groups.
  • Employ cross-tabs for comparing segment behaviors.
  • Use trend analysis over time for strategic decisions.
  • Always check for missing data and standardize features for clustering.
cross_tab = pd.crosstab(df['SpendingSegment'], df['Churn'])
print(cross_tab)
Churn              No  Yes
SpendingSegment           
Low              1624  214
Medium           1790  749
High             1760  906
segment_trend = df.groupby(['SpendingSegment','Contract']).size().unstack()
segment_trend.plot(kind='bar', stacked=True)
plt.title('Contract Types by Spending Segment')
plt.xlabel('Spending Segment')
plt.ylabel('Number of Customers')
plt.tight_layout()
plt.show()
No description has been provided for this image

End-to-end Market Research Problem: Segmenting and targeting customers#

  • Let's walk through the process from data to actionable insight.
  • We will identify key segments and recommend a strategy.
segment_counts = df.groupby(['Cluster','SpendingSegment']).size().unstack(fill_value=0)
print(segment_counts)
SpendingSegment   Low  Medium  High
Cluster                            
0                   0    1606  1075
1                   0     610  1591
2                1838     323     0
high_value = df[(df['SpendingSegment']=='High') & (df['Churn']=='No')]
print('Number of loyal high-value customers:', high_value.shape[0])
Number of loyal high-value customers: 1760
 

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.