Lesson 26 · Market Research Analytics in Python
Introduction to Market Segmentation: Essential Concepts for Market Research Analytics in Python
In this lesson, we will solve real-world market research problems by segmenting customers using real survey and behavioral data. Market segmentation helps…
- CourseMarket Research Analytics in Python
- Lesson26 of 56
- Video20 min
- FormatJupyter notebook · 18 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbIntroduction to Market Segmentation#
- In this lesson, we will solve real-world market research problems by segmenting customers using real survey and behavioral data.
- Market segmentation helps businesses understand different groups of customers so marketing, product, and service strategies can be tailored.
- You will learn to explore survey and customer data, identify useful segments, and generate meaningful business insights.
- By the end, you will be able to segment a customer base with Python and interpret segment differences for business decision-making.
import pandas as pd
import numpy as np
import openml
import matplotlib.pyplot as plt
import seaborn as sns
import warnings
warnings.filterwarnings('ignore')
What is market segmentation?#
- Market segmentation divides customers into distinct groups based on shared characteristics.
- Groups might be based on demographics, behaviors, survey responses or attitudes.
- Each segment should consist of customers expected to respond similarly to marketing and experience the product in a similar way.
- Market research data often comes from surveys or transaction logs.
- Responses may be numbers, multiple-choice answers, or text feedback.
- A common beginner mistake is assuming all customers are similar or treating categorical responses like numbers.
# Beginner Example 1: Load a real customer satisfaction survey dataset
dataset = openml.datasets.get_dataset(42178)
df_survey, _, _, _ = dataset.get_data(dataset_format='dataframe')
print(df_survey.shape)
print(df_survey.head(3))
# Beginner Example 2: Check for missing values
print(df_survey.isnull().sum())
# Beginner Example 3: Explore basic demographics in the survey
print(df_survey['gender'].value_counts())
print(df_survey['SeniorCitizen'].value_counts())
# Intermediate Example 1: Segment customers by contract type
contract_counts = df_survey['Contract'].value_counts()
print(contract_counts)
sns.barplot(y=contract_counts.index, x=contract_counts.values, orient='h')
plt.title('Customer Count by Contract Type')
plt.xlabel('Number of Customers')
plt.ylabel('Contract Type')
plt.show()
# Intermediate Example 2: Segment by Internet Service and Churn Rate
grouped = df_survey.groupby('InternetService')['Churn'].value_counts(normalize=True).unstack().fillna(0)
print(grouped)
grouped.plot(kind='bar', stacked=True)
plt.title('Churn Rate by Internet Service Type')
plt.ylabel('Proportion of Customers')
plt.xlabel('Internet Service Type')
plt.legend(title='Churn')
plt.show()
# Intermediate Example 3: Median tenure by segment (demographic)
median_tenure = df_survey.groupby('gender')['tenure'].median()
print(median_tenure)
# Advanced Example 1: Multi-way segmentation using multiple survey features
segments = df_survey.groupby(['Contract','gender']).size().unstack()
print(segments)
segments.plot(kind='bar', stacked=True)
plt.title('Customer Count by Contract Type and Gender')
plt.ylabel('Number of Customers')
plt.xlabel('Contract Type')
plt.show()
# Advanced Example 2: Identify high-value segments by spend
spend_by_segment = df_survey.groupby('InternetService')['MonthlyCharges'].mean().sort_values(ascending=False)
print(spend_by_segment)
sns.barplot(y=spend_by_segment.index, x=spend_by_segment.values, orient='h')
plt.title('Average Monthly Charges by Internet Service')
plt.xlabel('Average Monthly Charges ($)')
plt.ylabel('Internet Service')
plt.show()
# Advanced Example 3: Using Net Promoter Score (NPS) to segment customer loyalty
np.random.seed(42)
df_nps = pd.DataFrame({'CustomerID': range(1,501),
'Age': np.random.randint(18,70,500),
'Region': np.random.choice(['North','South','East','West'],500),
'NPS_Score': np.random.randint(0,11,500)})
df_nps['Segment'] = pd.cut(df_nps['NPS_Score'], bins=[-1,6,8,10], labels=['Detractor','Passive','Promoter'])
segment_counts = df_nps['Segment'].value_counts()
print(segment_counts)
sns.barplot(x=segment_counts.index, y=segment_counts.values)
plt.title('NPS Segmentation: Detractors, Passives, Promoters')
plt.xlabel('NPS Segment')
plt.ylabel('Number of Customers')
plt.show()
# Error Handling Example 1: What if a key grouping column is missing?
try:
print(df_survey.groupby('NonExistentColumn').size())
except Exception as e:
print('Error:', e)
# Error Handling Example 2: Handling missing survey responses
df_missing = df_survey.copy()
df_missing.loc[0, 'gender'] = np.nan
missing_count = df_missing['gender'].isnull().sum()
print('Missing gender responses:', missing_count)
filled = df_missing['gender'].fillna('Unknown')
print(filled.head(3))
# Error Handling Example 3: Watch out for misinterpreting Likert scales and NPS scores
avg_score = df_nps['NPS_Score'].mean()
print('Average NPS Score is:', avg_score)
if avg_score < 7:
print('Caution: Many business leaders misinterpret raw average NPS score. Use segment proportions for true insight.')
Analytics patterns and best practices#
- Always define business-relevant segments before deep analysis.
- Use cross-tabulation to compare groups on key performance indicators (e.g. NPS, spend, churn rate).
- Summarize categorical data by proportion, not just count.
- If building indices or scores, document how the index is constructed and why thresholds are chosen.
- Use trend charts to spot important shifts in segment behavior over time.
- Avoid treating survey codes or scaled responses as continuous variables without business logic.
# Analytics pattern: Cross-tabulation
crosstab = pd.crosstab(df_survey['gender'], df_survey['Churn'], normalize='index')
print(crosstab)
# Analytics pattern: Index and score construction for survey scales
def nps_index(df, score_col):
total = df.shape[0]
pct_promoters = np.mean(df[score_col] >= 9)
pct_detractors = np.mean(df[score_col] <= 6)
nps = (pct_promoters - pct_detractors) * 100
return round(nps,1)
print('Overall NPS Score:', nps_index(df_nps, 'NPS_Score'))
# Analytics pattern: Trend analysis of customer cohort behavior (synthetic cohort data setup)
np.random.seed(0)
dates = pd.date_range('2021-01-01', periods=24, freq='ME')
df_cohort = pd.DataFrame({'CustomerID': np.random.randint(1000,2000,len(dates)),
'Signup_Month': dates,
'Active_Users': np.random.randint(50,300,len(dates))})
plt.plot(df_cohort['Signup_Month'], df_cohort['Active_Users'], marker='o')
plt.title('Active Users by Signup Month (Trend)')
plt.ylabel('Number of Active Users')
plt.xlabel('Signup Month')
plt.xticks(rotation=45)
plt.tight_layout()
plt.show()
# End-to-end Example: From survey to insightWhich customer segments are at highest risk of churn?
segmentation = df_survey.groupby(['Contract','SeniorCitizen'])['Churn'].value_counts(normalize=True).unstack().fillna(0)
print(segmentation)
segmentation.plot(kind='bar', stacked=True)
plt.title('Churn Risk: By Contract and SeniorCitizen Status')
plt.ylabel('Proportion of Customers')
plt.xlabel('Contract and Senior Citizen Segment')
plt.legend(title='Churn')
plt.tight_layout()
plt.show()
high_risk = segmentation['Yes'].sort_values(ascending=False).head(3)
print('Top 3 segments at risk of churn:')
print(high_risk)
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



