Mathew K Analytics

Lesson 12 · Market Research Analytics in Python

Understanding Demographic Variables in Market Research Analytics with Python

In this lesson, we will learn how to analyze and understand demographic variables using real-world datasets. Understanding demographics is key to segmenting…

⬇ Download notebookOpen in Colab ↗

📓 Full notebook

Download .ipynb

Understanding Demographic Variables in Market Research#

  • In this lesson, we will learn how to analyze and understand demographic variables using real-world datasets.
  • Understanding demographics is key to segmenting customers and analyzing customer satisfaction.
  • Businesses often use demographic analysis to target marketing and improve customer experiences.
  • You will use Python to summarize, visualize, and draw insights from customer surveys and demographic data.
import pandas as pd
import numpy as np
import openml
import matplotlib.pyplot as plt
import seaborn as sns
import warnings
warnings.filterwarnings('ignore')

What Are Demographic Variables?#

  • Demographic variables describe key characteristics about people (like age, gender, income).
  • These variables help us group, segment, and understand customer behaviors.
  • In survey and customer data, demographic variables are often columns like 'age', 'gender', or 'region'.
  • Mistakes beginners make: ignoring missing values, misinterpreting categories, or failing to segment properly.

Exploring Customer Demographics: Beginner Example 1#

  • Let us start by examining the structure and columns of a real customer satisfaction dataset.
dataset = openml.datasets.get_dataset(42178)
df, _, _, _ = dataset.get_data(dataset_format='dataframe')
print(df.shape)
print(df[['gender','SeniorCitizen','Partner','Dependents','tenure']].head())
(7043, 20)
   gender  SeniorCitizen Partner Dependents  tenure
0  Female              0     Yes         No       1
1    Male              0      No         No      34
2    Male              0      No         No       2
3    Male              0      No         No      45
4  Female              0      No         No       2

Beginner Example 2: Visualizing Gender and Senior Status#

  • Let us create a simple countplot to visualize counts of gender and senior citizenship.
sns.countplot(data=df, x='gender', hue='SeniorCitizen')
plt.title('Customer Count by Gender and Senior Status')
plt.show()
No description has been provided for this image

Beginner Example 3: Summarizing Numeric Demographic Variables#

  • We will look at basic statistics for tenure (how long a customer has stayed).
  • This helps us understand the distribution and identify any unusual values.
print(df['tenure'].describe())
count    7043.000000
mean       32.371149
std        24.559481
min         0.000000
25%         9.000000
50%        29.000000
75%        55.000000
max        72.000000
Name: tenure, dtype: float64

Intermediate Example 1: Cross-Tabulating Demographics and Churn#

  • We can see how customer demographic groups relate to churn (customer loss).
  • Cross-tabulation lets us compare demographic segments' churn rates.
churn_crosstab = pd.crosstab(df['SeniorCitizen'], df['Churn'], normalize='index')
print(churn_crosstab)
Churn                No       Yes
SeniorCitizen                    
0              0.763938  0.236062
1              0.583187  0.416813

Intermediate Example 2: Exploring Demographics in Marketing Campaign Data#

  • Let us switch to a marketing campaign dataset and examine its demographic structure.
dataset2 = openml.datasets.get_dataset(1461)
df2, _, _, _ = dataset2.get_data(dataset_format='dataframe')
df2.columns = ['age','job','marital','education','default','balance','housing','loan','contact','day','month','duration','campaign','pdays','previous','poutcome','response']
print(df2[['age','job','marital','education']].head())
   age           job  marital  education
0   58    management  married   tertiary
1   44    technician   single  secondary
2   33  entrepreneur  married  secondary
3   47   blue-collar  married    unknown
4   33       unknown   single    unknown

Intermediate Example 3: Demographic Profiles of Responders#

  • We will summarize the ages of customers who responded to the marketing campaign.
responders = df2[df2['response'] == '2']
print(responders['age'].describe())
count    5289.000000
mean       41.670070
std        13.497781
min        18.000000
25%        31.000000
50%        38.000000
75%        50.000000
max        95.000000
Name: age, dtype: float64

Advanced Example 1: Segmenting Customers by Multiple Demographics#

  • We can use groupby to segment the marketing data by both education and marital status.
  • This helps us find high-potential customer groups for targeted advertising.
grouped = df2.groupby(['education', 'marital']).size().unstack()
print(grouped)
marital    divorced  married  single
education                           
primary         752     5246     853
secondary      2815    13770    6617
tertiary       1471     7038    4792
unknown         169     1160     528

Advanced Example 2: Visualizing Churn Rate by Age Groups#

  • Now let us create age bins and plot churn rates for different age groups, using the customer satisfaction dataset.
df['age_group'] = pd.cut(df['tenure'], bins=[0,12,24,48,72], labels=['<1yr', '1-2yr', '2-4yr', '4-6yr'])
churn_by_age = df.groupby('age_group')['Churn'].value_counts(normalize=True).unstack().fillna(0)
churn_by_age['Churn_Rate'] = churn_by_age['Yes']
churn_by_age['Churn_Rate'].plot(kind='bar', color='tomato', title='Churn Rate by Customer Tenure Group')
plt.ylabel('Churn Rate')
plt.show()
No description has been provided for this image

Advanced Example 3: Comparing Average NPS Scores by Region#

  • Let us now look at a synthetic Net Promoter Score survey dataset and compare NPS by region.
np.random.seed(42)
df_nps = pd.DataFrame({'CustomerID': range(1,501),
                      'Age': np.random.randint(18,70,500),
                      'Region': np.random.choice(['North','South','East','West'],500),
                      'NPS_Score': np.random.randint(0,11,500)})
region_nps = df_nps.groupby('Region')['NPS_Score'].mean()
print(region_nps)
Region
East     4.504425
North    5.214765
South    4.719008
West     5.025641
Name: NPS_Score, dtype: float64

Error Handling: Missing Survey Responses#

  • Real data often include missing demographic values.
  • Let us explore how to detect and handle such issues for robust market research analyses.
# Introduce missingness for demonstration
df_miss = df.copy()
df_miss.loc[df_miss.sample(frac=0.05, random_state=42).index, 'gender'] = np.nan
print(df_miss['gender'].isna().sum())
352

Debugging Example: Incorrect Grouping#

  • Mistakes in groupby or crosstab can give misleading results.
  • We will intentionally group with the wrong key and explore the impact.
# Wrong grouping: grouping by random index
df2['group'] = np.random.choice(['A','B'], size=len(df2), replace=True)
wrong_grouping = df2.groupby('group')['response'].value_counts(normalize=True).unstack()
print(wrong_grouping)
response         1         2
group                       
A         0.882158  0.117842
B         0.883873  0.116127

Debugging Example: Misinterpreting NPS Scores#

  • Beginners sometimes treat NPS scores as continuous, but they are an index (0-10).
  • Let us see the spread and think about what it means.
plt.hist(df_nps['NPS_Score'], bins=np.arange(12)-0.5, color='steelblue', edgecolor='black')
plt.xticks(range(0,11))
plt.title('Distribution of NPS Scores (0-10)')
plt.xlabel('NPS Score')
plt.ylabel('Number of Respondents')
plt.show()
No description has been provided for this image

Market Research Best Practices: Segmentation and Cross-Tabulation#

  • Segment your data before analysis: group by demographics for targeted insights.
  • Cross-tabulation compares outcomes across groups (like churn by gender or age).
  • Always check for missing or inconsistent demographic values first.

Market Research Analytics Patterns#

  • Construct indices and scores only when well-defined (like NPS or customer satisfaction index).
  • Visualize trends over time (tenure, signup month, repeat purchases) to uncover hidden insights.
  • Use practice exercises to reinforce learning: e.g., plot churn by education level.

End-to-End Market Research Example: From Raw Data to Insight#

  • Now let us start with marketing campaign data and come up with a segmentation-based recommendation.
  • We will:
    1. Examine demographics and response rates.
    1. Identify the segment with the highest response ratio.
    1. Recommend targeting this segment in future campaigns.
edu_response = pd.crosstab(df2['education'], df2['response'], normalize='index')
print(edu_response)
top_edu = edu_response['2'].idxmax()
top_rate = edu_response['2'].max()
print(f"Highest response rate: {top_edu} ({top_rate:.2%})")
response          1         2
education                    
primary    0.913735  0.086265
secondary  0.894406  0.105594
tertiary   0.849936  0.150064
unknown    0.864297  0.135703
Highest response rate: tertiary (15.01%)
 

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.