Mathew K Analytics

Lesson 26 · Market Research Analytics in Python

Introduction to Market Segmentation: Essential Concepts for Market Research Analytics in Python

In this lesson, we will solve real-world market research problems by segmenting customers using real survey and behavioral data. Market segmentation helps…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Introduction to Market Segmentation#

  • In this lesson, we will solve real-world market research problems by segmenting customers using real survey and behavioral data.
  • Market segmentation helps businesses understand different groups of customers so marketing, product, and service strategies can be tailored.
  • You will learn to explore survey and customer data, identify useful segments, and generate meaningful business insights.
  • By the end, you will be able to segment a customer base with Python and interpret segment differences for business decision-making.
import pandas as pd
import numpy as np
import openml
import matplotlib.pyplot as plt
import seaborn as sns
import warnings
warnings.filterwarnings('ignore')

What is market segmentation?#

  • Market segmentation divides customers into distinct groups based on shared characteristics.
  • Groups might be based on demographics, behaviors, survey responses or attitudes.
  • Each segment should consist of customers expected to respond similarly to marketing and experience the product in a similar way.
  • Market research data often comes from surveys or transaction logs.
  • Responses may be numbers, multiple-choice answers, or text feedback.
  • A common beginner mistake is assuming all customers are similar or treating categorical responses like numbers.
# Beginner Example 1: Load a real customer satisfaction survey dataset
dataset = openml.datasets.get_dataset(42178)
df_survey, _, _, _ = dataset.get_data(dataset_format='dataframe')
print(df_survey.shape)
print(df_survey.head(3))
(7043, 20)
   gender  SeniorCitizen Partner Dependents  tenure PhoneService  \
0  Female              0     Yes         No       1           No   
1    Male              0      No         No      34          Yes   
2    Male              0      No         No       2          Yes   

      MultipleLines InternetService OnlineSecurity OnlineBackup  \
0  No phone service             DSL             No          Yes   
1                No             DSL            Yes           No   
2                No             DSL            Yes          Yes   

  DeviceProtection TechSupport StreamingTV StreamingMovies        Contract  \
0               No          No          No              No  Month-to-month   
1              Yes          No          No              No        One year   
2               No          No          No              No  Month-to-month   

  PaperlessBilling     PaymentMethod  MonthlyCharges TotalCharges Churn  
0              Yes  Electronic check           29.85        29.85    No  
1               No      Mailed check           56.95       1889.5    No  
2              Yes      Mailed check           53.85       108.15   Yes  
# Beginner Example 2: Check for missing values
print(df_survey.isnull().sum())
gender              0
SeniorCitizen       0
Partner             0
Dependents          0
tenure              0
PhoneService        0
MultipleLines       0
InternetService     0
OnlineSecurity      0
OnlineBackup        0
DeviceProtection    0
TechSupport         0
StreamingTV         0
StreamingMovies     0
Contract            0
PaperlessBilling    0
PaymentMethod       0
MonthlyCharges      0
TotalCharges        0
Churn               0
dtype: int64
# Beginner Example 3: Explore basic demographics in the survey
print(df_survey['gender'].value_counts())
print(df_survey['SeniorCitizen'].value_counts())
gender
Male      3555
Female    3488
Name: count, dtype: int64
SeniorCitizen
0    5901
1    1142
Name: count, dtype: int64
# Intermediate Example 1: Segment customers by contract type
contract_counts = df_survey['Contract'].value_counts()
print(contract_counts)
sns.barplot(y=contract_counts.index, x=contract_counts.values, orient='h')
plt.title('Customer Count by Contract Type')
plt.xlabel('Number of Customers')
plt.ylabel('Contract Type')
plt.show()
Contract
Month-to-month    3875
Two year          1695
One year          1473
Name: count, dtype: int64
No description has been provided for this image
# Intermediate Example 2: Segment by Internet Service and Churn Rate
grouped = df_survey.groupby('InternetService')['Churn'].value_counts(normalize=True).unstack().fillna(0)
print(grouped)
grouped.plot(kind='bar', stacked=True)
plt.title('Churn Rate by Internet Service Type')
plt.ylabel('Proportion of Customers')
plt.xlabel('Internet Service Type')
plt.legend(title='Churn')
plt.show()
Churn                  No       Yes
InternetService                    
DSL              0.810409  0.189591
Fiber optic      0.581072  0.418928
No               0.925950  0.074050
No description has been provided for this image
# Intermediate Example 3: Median tenure by segment (demographic)
median_tenure = df_survey.groupby('gender')['tenure'].median()
print(median_tenure)
gender
Female    29.0
Male      29.0
Name: tenure, dtype: float64
# Advanced Example 1: Multi-way segmentation using multiple survey features
segments = df_survey.groupby(['Contract','gender']).size().unstack()
print(segments)
segments.plot(kind='bar', stacked=True)
plt.title('Customer Count by Contract Type and Gender')
plt.ylabel('Number of Customers')
plt.xlabel('Contract Type')
plt.show()
gender          Female  Male
Contract                    
Month-to-month    1925  1950
One year           718   755
Two year           845   850
No description has been provided for this image
# Advanced Example 2: Identify high-value segments by spend
spend_by_segment = df_survey.groupby('InternetService')['MonthlyCharges'].mean().sort_values(ascending=False)
print(spend_by_segment)
sns.barplot(y=spend_by_segment.index, x=spend_by_segment.values, orient='h')
plt.title('Average Monthly Charges by Internet Service')
plt.xlabel('Average Monthly Charges ($)')
plt.ylabel('Internet Service')
plt.show()
InternetService
Fiber optic    91.500129
DSL            58.102169
No             21.079194
Name: MonthlyCharges, dtype: float64
No description has been provided for this image
# Advanced Example 3: Using Net Promoter Score (NPS) to segment customer loyalty
np.random.seed(42)
df_nps = pd.DataFrame({'CustomerID': range(1,501),
                      'Age': np.random.randint(18,70,500),
                      'Region': np.random.choice(['North','South','East','West'],500),
                      'NPS_Score': np.random.randint(0,11,500)})
df_nps['Segment'] = pd.cut(df_nps['NPS_Score'], bins=[-1,6,8,10], labels=['Detractor','Passive','Promoter'])
segment_counts = df_nps['Segment'].value_counts()
print(segment_counts)
sns.barplot(x=segment_counts.index, y=segment_counts.values)
plt.title('NPS Segmentation: Detractors, Passives, Promoters')
plt.xlabel('NPS Segment')
plt.ylabel('Number of Customers')
plt.show()
Segment
Detractor    323
Passive       92
Promoter      85
Name: count, dtype: int64
No description has been provided for this image
# Error Handling Example 1: What if a key grouping column is missing?
try:
    print(df_survey.groupby('NonExistentColumn').size())
except Exception as e:
    print('Error:', e)
Error: 'NonExistentColumn'
# Error Handling Example 2: Handling missing survey responses
df_missing = df_survey.copy()
df_missing.loc[0, 'gender'] = np.nan
missing_count = df_missing['gender'].isnull().sum()
print('Missing gender responses:', missing_count)
filled = df_missing['gender'].fillna('Unknown')
print(filled.head(3))
Missing gender responses: 1
0    Unknown
1       Male
2       Male
Name: gender, dtype: object
# Error Handling Example 3: Watch out for misinterpreting Likert scales and NPS scores
avg_score = df_nps['NPS_Score'].mean()
print('Average NPS Score is:', avg_score)
if avg_score < 7:
    print('Caution: Many business leaders misinterpret raw average NPS score. Use segment proportions for true insight.')
Average NPS Score is: 4.89
Caution: Many business leaders misinterpret raw average NPS score. Use segment proportions for true insight.

Analytics patterns and best practices#

  • Always define business-relevant segments before deep analysis.
  • Use cross-tabulation to compare groups on key performance indicators (e.g. NPS, spend, churn rate).
  • Summarize categorical data by proportion, not just count.
  • If building indices or scores, document how the index is constructed and why thresholds are chosen.
  • Use trend charts to spot important shifts in segment behavior over time.
  • Avoid treating survey codes or scaled responses as continuous variables without business logic.
# Analytics pattern: Cross-tabulation
crosstab = pd.crosstab(df_survey['gender'], df_survey['Churn'], normalize='index')
print(crosstab)
Churn         No       Yes
gender                    
Female  0.730791  0.269209
Male    0.738397  0.261603
# Analytics pattern: Index and score construction for survey scales
def nps_index(df, score_col):
    total = df.shape[0]
    pct_promoters = np.mean(df[score_col] >= 9)
    pct_detractors = np.mean(df[score_col] <= 6)
    nps = (pct_promoters - pct_detractors) * 100
    return round(nps,1)
print('Overall NPS Score:', nps_index(df_nps, 'NPS_Score'))
Overall NPS Score: -47.6
# Analytics pattern: Trend analysis of customer cohort behavior (synthetic cohort data setup)
np.random.seed(0)
dates = pd.date_range('2021-01-01', periods=24, freq='ME')
df_cohort = pd.DataFrame({'CustomerID': np.random.randint(1000,2000,len(dates)),
                         'Signup_Month': dates,
                         'Active_Users': np.random.randint(50,300,len(dates))})
plt.plot(df_cohort['Signup_Month'], df_cohort['Active_Users'], marker='o')
plt.title('Active Users by Signup Month (Trend)')
plt.ylabel('Number of Active Users')
plt.xlabel('Signup Month')
plt.xticks(rotation=45)
plt.tight_layout()
plt.show()
No description has been provided for this image
# End-to-end Example: From survey to insightWhich customer segments are at highest risk of churn?
segmentation = df_survey.groupby(['Contract','SeniorCitizen'])['Churn'].value_counts(normalize=True).unstack().fillna(0)
print(segmentation)
segmentation.plot(kind='bar', stacked=True)
plt.title('Churn Risk: By Contract and SeniorCitizen Status')
plt.ylabel('Proportion of Customers')
plt.xlabel('Contract and Senior Citizen Segment')
plt.legend(title='Churn')
plt.tight_layout()
plt.show()
high_risk = segmentation['Yes'].sort_values(ascending=False).head(3)
print('Top 3 segments at risk of churn:')
print(high_risk)
Churn                               No       Yes
Contract       SeniorCitizen                    
Month-to-month 0              0.604302  0.395698
               1              0.453532  0.546468
One year       0              0.893219  0.106781
               1              0.847368  0.152632
Two year       0              0.972903  0.027097
               1              0.958621  0.041379
No description has been provided for this image
Top 3 segments at risk of churn:
Contract        SeniorCitizen
Month-to-month  1                0.546468
                0                0.395698
One year        1                0.152632
Name: Yes, dtype: float64
 

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.