Lesson 27 · Market Research Analytics in Python
Demographic Segmentation Analysis for Market Research in Python
We will learn how to divide customers into groups based on demographic attributes. This helps businesses understand who their key customers are and how…
- CourseMarket Research Analytics in Python
- Lesson27 of 56
- Video21 min
- FormatJupyter notebook · 27 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbDemographic Segmentation Analysis in Customer Analytics#
- We will learn how to divide customers into groups based on demographic attributes.
- This helps businesses understand who their key customers are and how needs differ.
- You will use real-world datasets to uncover patterns by age group, gender, region, and more.
- Insights can guide marketing, product development, and customer retention strategies.
import pandas as pd
import numpy as np
import openml
import warnings
warnings.filterwarnings('ignore')
Understanding demographic segmentation and the datasets#
- Demographic segmentation means dividing customers based on attributes such as age, gender, income, region, or education.
- Real business survey or customer datasets usually include categorical and numeric columns.
- Columns may represent gender (M/F), age group, region, and responses to satisfaction or NPS questions.
- Beginners often forget to check for missing data, or may group by the wrong column.
- Some mistakes include mixing numerical and categorical features in aggregation, or mislabeling demographic buckets.
dataset = openml.datasets.get_dataset(42178)
df, _, _, _ = dataset.get_data(dataset_format='dataframe')
print(df.shape)
print(df.head(3))
gender_counts = df['gender'].value_counts()
print(gender_counts)
senior_pct = df['SeniorCitizen'].value_counts(normalize=True) * 100
print(senior_pct.round(2))
partner_counts = df['Partner'].value_counts()
print(partner_counts)
dependents_counts = df['Dependents'].value_counts()
print(dependents_counts)
age_groups = pd.cut(df['tenure'], bins=[0, 12, 36, 72], labels=['<1yr', '1-3yrs', '3-6yrs'])
df['TenureGroup'] = age_groups
grouped_tenure = df.groupby('TenureGroup').size()
print(grouped_tenure)
churn_by_gender = df.groupby('gender')['Churn'].value_counts().unstack().fillna(0)
print(churn_by_gender)
avg_monthly_by_senior = df.groupby('SeniorCitizen')['MonthlyCharges'].mean()
print(avg_monthly_by_senior.round(2))
contract_churn = df.groupby('Contract')['Churn'].value_counts(normalize=True).unstack().fillna(0) * 100
print(contract_churn.round(1))
internet_by_senior = pd.crosstab(df['SeniorCitizen'], df['InternetService'])
print(internet_by_senior)
churn_family = df.groupby(['Partner', 'Dependents'])['Churn'].value_counts(normalize=True).unstack().fillna(0) * 100
print(churn_family.round(2))
dataset2 = openml.datasets.get_dataset(1461)
df2, _, _, _ = dataset2.get_data(dataset_format='dataframe')
df2.columns = ['age','job','marital','education','default','balance','housing','loan','contact','day','month','duration','campaign','pdays','previous','poutcome','response']
age_brackets = pd.cut(df2['age'], bins=[15,25,35,50,100], labels=['16-25','26-35','36-50','51+'])
segment_counts = age_brackets.value_counts()
print(segment_counts)
crosstab = pd.crosstab([df2['education'], age_brackets], df2['response'])
print(crosstab)
df2['balance_cat'] = pd.cut(df2['balance'], bins=[-np.inf,0,1000,5000,np.inf], labels=['Debt','Low','Medium','High'])
response_by_segment = df2.groupby(['job','balance_cat'])['response'].value_counts(normalize=True).unstack().fillna(0) * 100
print(response_by_segment.head(6).round(1))
missing_ages = df['tenure'].isnull().sum()
print(f'Missing tenure entries: {missing_ages}')
if 'region' in df.columns:
print(df['region'].value_counts())
else:
print('No region field available in this dataset.')
# Example: Incorrect groupings due to a typo
try:
wrong_seg = df.groupby('Gennder').size()
except Exception as e:
print(f'Error: {e}')
n_na = df['MonthlyCharges'].isnull().sum()
if n_na > 0:
print(f'Found {n_na} missing values in MonthlyCharges. Consider filling with median.')
df['MonthlyCharges'] = df['MonthlyCharges'].fillna(df['MonthlyCharges'].median())
else:
print('No missing values detected.')
# Example: Misinterpretation of NPS-like score
if 'Churn' in df.columns:
nps_proxy = (df['Churn']=='No').mean() * 100
print(f'Percent customers satisfied / loyal (proxy): {nps_proxy:.2f}%')
else:
print('NPS field unavailable; use a satisfied/loyalty proxy.')
cross_demo = pd.crosstab(df['gender'], df['SeniorCitizen'], margins=True)
print(cross_demo)
df['FamilySegment'] = (df['Partner']=='Yes').astype(str) + '_' + (df['Dependents']=='Yes').astype(str)
result = df.groupby('FamilySegment')['MonthlyCharges'].mean()
print(result)
churn_index = df.groupby('Contract')['Churn'].apply(lambda x: (x=='Yes').mean()*100)
print(churn_index.round(2))
trends = df.groupby('SeniorCitizen')['tenure'].mean()
print(trends.round(1))
np.random.seed(42)
edemo = pd.DataFrame({
'CustomerID': range(1,201),
'Gender': np.random.choice(['Male','Female'],200),
'AgeGroup': np.random.choice(['18-25','26-35','36-50','51+'],200,p=[0.2,0.3,0.3,0.2]),
'NPS': np.random.randint(0,11,200)
})
print(edemo.head(3))
nps_segment = edemo.groupby(['Gender','AgeGroup'])['NPS'].mean().unstack()
print(nps_segment.round(2))
# Business insight: flag weak segments for attention
min_nps = nps_segment.min().min()
if min_nps < 5:
print('ALERT: A demographic segment has a low NPS!')
else:
print('All segments are reasonably satisfied.')
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



