Mathew K Analytics

Lesson 27 · Market Research Analytics in Python

Demographic Segmentation Analysis for Market Research in Python

We will learn how to divide customers into groups based on demographic attributes. This helps businesses understand who their key customers are and how…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Demographic Segmentation Analysis in Customer Analytics#

  • We will learn how to divide customers into groups based on demographic attributes.
  • This helps businesses understand who their key customers are and how needs differ.
  • You will use real-world datasets to uncover patterns by age group, gender, region, and more.
  • Insights can guide marketing, product development, and customer retention strategies.
import pandas as pd
import numpy as np
import openml
import warnings
warnings.filterwarnings('ignore')

Understanding demographic segmentation and the datasets#

  • Demographic segmentation means dividing customers based on attributes such as age, gender, income, region, or education.
  • Real business survey or customer datasets usually include categorical and numeric columns.
  • Columns may represent gender (M/F), age group, region, and responses to satisfaction or NPS questions.
  • Beginners often forget to check for missing data, or may group by the wrong column.
  • Some mistakes include mixing numerical and categorical features in aggregation, or mislabeling demographic buckets.
dataset = openml.datasets.get_dataset(42178)
df, _, _, _ = dataset.get_data(dataset_format='dataframe')
print(df.shape)
print(df.head(3))
(7043, 20)
   gender  SeniorCitizen Partner Dependents  tenure PhoneService  \
0  Female              0     Yes         No       1           No   
1    Male              0      No         No      34          Yes   
2    Male              0      No         No       2          Yes   

      MultipleLines InternetService OnlineSecurity OnlineBackup  \
0  No phone service             DSL             No          Yes   
1                No             DSL            Yes           No   
2                No             DSL            Yes          Yes   

  DeviceProtection TechSupport StreamingTV StreamingMovies        Contract  \
0               No          No          No              No  Month-to-month   
1              Yes          No          No              No        One year   
2               No          No          No              No  Month-to-month   

  PaperlessBilling     PaymentMethod  MonthlyCharges TotalCharges Churn  
0              Yes  Electronic check           29.85        29.85    No  
1               No      Mailed check           56.95       1889.5    No  
2              Yes      Mailed check           53.85       108.15   Yes  
gender_counts = df['gender'].value_counts()
print(gender_counts)
gender
Male      3555
Female    3488
Name: count, dtype: int64
senior_pct = df['SeniorCitizen'].value_counts(normalize=True) * 100
print(senior_pct.round(2))
SeniorCitizen
0    83.79
1    16.21
Name: proportion, dtype: float64
partner_counts = df['Partner'].value_counts()
print(partner_counts)
dependents_counts = df['Dependents'].value_counts()
print(dependents_counts)
Partner
No     3641
Yes    3402
Name: count, dtype: int64
Dependents
No     4933
Yes    2110
Name: count, dtype: int64
age_groups = pd.cut(df['tenure'], bins=[0, 12, 36, 72], labels=['<1yr', '1-3yrs', '3-6yrs'])
df['TenureGroup'] = age_groups
grouped_tenure = df.groupby('TenureGroup').size()
print(grouped_tenure)
TenureGroup
<1yr      2175
1-3yrs    1856
3-6yrs    3001
dtype: int64
churn_by_gender = df.groupby('gender')['Churn'].value_counts().unstack().fillna(0)
print(churn_by_gender)
Churn     No  Yes
gender           
Female  2549  939
Male    2625  930
avg_monthly_by_senior = df.groupby('SeniorCitizen')['MonthlyCharges'].mean()
print(avg_monthly_by_senior.round(2))
SeniorCitizen
0    61.85
1    79.82
Name: MonthlyCharges, dtype: float64
contract_churn = df.groupby('Contract')['Churn'].value_counts(normalize=True).unstack().fillna(0) * 100
print(contract_churn.round(1))
Churn             No   Yes
Contract                  
Month-to-month  57.3  42.7
One year        88.7  11.3
Two year        97.2   2.8
internet_by_senior = pd.crosstab(df['SeniorCitizen'], df['InternetService'])
print(internet_by_senior)
InternetService   DSL  Fiber optic    No
SeniorCitizen                           
0                2162         2265  1474
1                 259          831    52
churn_family = df.groupby(['Partner', 'Dependents'])['Churn'].value_counts(normalize=True).unstack().fillna(0) * 100
print(churn_family.round(2))
Churn                  No    Yes
Partner Dependents              
No      No          65.76  34.24
        Yes         78.67  21.33
Yes     No          74.59  25.41
        Yes         85.76  14.24
dataset2 = openml.datasets.get_dataset(1461)
df2, _, _, _ = dataset2.get_data(dataset_format='dataframe')
df2.columns = ['age','job','marital','education','default','balance','housing','loan','contact','day','month','duration','campaign','pdays','previous','poutcome','response']
age_brackets = pd.cut(df2['age'], bins=[15,25,35,50,100], labels=['16-25','26-35','36-50','51+'])
segment_counts = age_brackets.value_counts()
print(segment_counts)
age
36-50    19049
26-35    15571
51+       9255
16-25     1336
Name: count, dtype: int64
crosstab = pd.crosstab([df2['education'], age_brackets], df2['response'])
print(crosstab)
response            1    2
education age             
primary   16-25    73   24
          26-35  1173   85
          36-50  2980  188
          51+    2034  294
secondary 16-25   693  196
          26-35  7517  854
          36-50  9029  838
          51+    3513  562
tertiary  16-25   179   67
          26-35  4703  887
          36-50  4541  676
          51+    1882  366
unknown   16-25    71   33
          26-35   309   43
          36-50   712   85
          51+     513   91
df2['balance_cat'] = pd.cut(df2['balance'], bins=[-np.inf,0,1000,5000,np.inf], labels=['Debt','Low','Medium','High'])
response_by_segment = df2.groupby(['job','balance_cat'])['response'].value_counts(normalize=True).unstack().fillna(0) * 100
print(response_by_segment.head(6).round(1))
response                    1     2
job         balance_cat            
admin.      Debt         92.3   7.7
            Low          87.9  12.1
            Medium       85.1  14.9
            High         84.6  15.4
blue-collar Debt         94.8   5.2
            Low          92.8   7.2
missing_ages = df['tenure'].isnull().sum()
print(f'Missing tenure entries: {missing_ages}')
Missing tenure entries: 0
if 'region' in df.columns:
    print(df['region'].value_counts())
else:
    print('No region field available in this dataset.')
No region field available in this dataset.
# Example: Incorrect groupings due to a typo
try:
    wrong_seg = df.groupby('Gennder').size()
except Exception as e:
    print(f'Error: {e}')
Error: 'Gennder'
n_na = df['MonthlyCharges'].isnull().sum()
if n_na > 0:
    print(f'Found {n_na} missing values in MonthlyCharges. Consider filling with median.')
    df['MonthlyCharges'] = df['MonthlyCharges'].fillna(df['MonthlyCharges'].median())
else:
    print('No missing values detected.')
No missing values detected.
# Example: Misinterpretation of NPS-like score
if 'Churn' in df.columns:
    nps_proxy = (df['Churn']=='No').mean() * 100
    print(f'Percent customers satisfied / loyal (proxy): {nps_proxy:.2f}%')
else:
    print('NPS field unavailable; use a satisfied/loyalty proxy.')
Percent customers satisfied / loyal (proxy): 73.46%
cross_demo = pd.crosstab(df['gender'], df['SeniorCitizen'], margins=True)
print(cross_demo)
SeniorCitizen     0     1   All
gender                         
Female         2920   568  3488
Male           2981   574  3555
All            5901  1142  7043
df['FamilySegment'] = (df['Partner']=='Yes').astype(str) + '_' + (df['Dependents']=='Yes').astype(str)
result = df.groupby('FamilySegment')['MonthlyCharges'].mean()
print(result)
FamilySegment
False_False    62.983735
False_True     52.507202
True_False     74.977737
True_True      60.970069
Name: MonthlyCharges, dtype: float64
churn_index = df.groupby('Contract')['Churn'].apply(lambda x: (x=='Yes').mean()*100)
print(churn_index.round(2))
Contract
Month-to-month    42.71
One year          11.27
Two year           2.83
Name: Churn, dtype: float64
trends = df.groupby('SeniorCitizen')['tenure'].mean()
print(trends.round(1))
SeniorCitizen
0    32.2
1    33.3
Name: tenure, dtype: float64
np.random.seed(42)
edemo = pd.DataFrame({
    'CustomerID': range(1,201),
    'Gender': np.random.choice(['Male','Female'],200),
    'AgeGroup': np.random.choice(['18-25','26-35','36-50','51+'],200,p=[0.2,0.3,0.3,0.2]),
    'NPS': np.random.randint(0,11,200)
})
print(edemo.head(3))
   CustomerID  Gender AgeGroup  NPS
0           1    Male    18-25    3
1           2  Female    36-50    2
2           3    Male    26-35    6
nps_segment = edemo.groupby(['Gender','AgeGroup'])['NPS'].mean().unstack()
print(nps_segment.round(2))
AgeGroup  18-25  26-35  36-50   51+
Gender                             
Female     5.18   5.23   4.03  5.23
Male       5.38   4.19   4.85  5.04
# Business insight: flag weak segments for attention
min_nps = nps_segment.min().min()
if min_nps < 5:
    print('ALERT: A demographic segment has a low NPS!')
else:
    print('All segments are reasonably satisfied.')
ALERT: A demographic segment has a low NPS!
 

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.