Mathew K Analytics

Lesson 34 · Market Research Analytics in Python

Chi-Square Tests for Analyzing Survey Relationships Using Python | Market Research Analytics

We will learn to analyze whether customer survey responses are related to key demographics or behaviors. This analysis helps businesses discover patterns…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Chi-Square Tests for Survey Relationships#

  • We will learn to analyze whether customer survey responses are related to key demographics or behaviors.
  • This analysis helps businesses discover patterns such as which customer groups are more likely to churn or recommend a service.
  • By applying chi-square tests, we will detect whether variables like satisfaction, churn, or NPS differ across segments.
  • The goal is to move from raw survey data to actionable insights for business improvements.
import warnings
warnings.filterwarnings('ignore')
import pandas as pd
import numpy as np
import openml
from scipy.stats import chi2_contingency
import matplotlib.pyplot as plt
import seaborn as sns

Core Concepts: What Are Chi-Square Tests in Market Research?#

  • Customer or market surveys collect responses (like satisfaction or NPS) along with demographic or behavioral data.
  • Chi-square tests help us check if two categorical variables (such as gender and churn) are statistically related.
  • Survey data may be structured with columns like gender, age group, product chosen, and survey outcome.
  • Common mistakes include using chi-square on continuous data or not checking group sample sizes.
  • Beginner analysts sometimes ignore data cleaning or the risk of empty cells in the data.
dataset = openml.datasets.get_dataset(42178)
df, _, _, _ = dataset.get_data(dataset_format='dataframe')
print(df.shape)
print(df[['gender', 'SeniorCitizen', 'Churn']].head())
(7043, 20)
   gender  SeniorCitizen Churn
0  Female              0    No
1    Male              0    No
2    Male              0   Yes
3    Male              0    No
4  Female              0   Yes
ct = pd.crosstab(df['gender'], df['Churn'])
print(ct)
Churn     No  Yes
gender           
Female  2549  939
Male    2625  930
chi2, p, dof, ex = chi2_contingency(ct)
print(f"Chi2 Statistic: {chi2:.3f}, p-value: {p:.4f}")
Chi2 Statistic: 0.484, p-value: 0.4866
sns.countplot(x='gender', hue='Churn', data=df)
plt.title('Churn by Gender')
plt.show()
No description has been provided for this image
dataset = openml.datasets.get_dataset(1461)
df_marketing, _, _, _ = dataset.get_data(dataset_format='dataframe')
df_marketing.columns = ['age','job','marital','education','default','balance','housing','loan','contact','day','month','duration','campaign','pdays','previous','poutcome','response']
print(df_marketing[['job','response']].head())
            job response
0    management        1
1    technician        1
2  entrepreneur        1
3   blue-collar        1
4       unknown        1
ct_job = pd.crosstab(df_marketing['job'], df_marketing['response'])
print(ct_job.head())
response         1     2
job                     
admin.        4540   631
blue-collar   9024   708
entrepreneur  1364   123
housemaid     1131   109
management    8157  1301
chi2, p, dof, ex = chi2_contingency(ct_job)
print(f"Chi2 Statistic: {chi2:.2f}, p-value: {p:.5f}")
if p < 0.05:
    print('There is a statistically significant relationship between job type and campaign response.')
else:
    print('No significant relationship between job type and campaign response.')
Chi2 Statistic: 836.11, p-value: 0.00000
There is a statistically significant relationship between job type and campaign response.
ct_edu = pd.crosstab(df_marketing['education'], df_marketing['response'])
chi2, p, dof, ex = chi2_contingency(ct_edu)
print(ct_edu)
print(f"Chi2: {chi2:.2f}, p-value: {p:.4f}")
response       1     2
education             
primary     6260   591
secondary  20752  2450
tertiary   11305  1996
unknown     1605   252
Chi2: 238.92, p-value: 0.0000
np.random.seed(42)
df_nps = pd.DataFrame({'CustomerID': range(1,501),
                      'Age': np.random.randint(18,70,500),
                      'Region': np.random.choice(['North','South','East','West'],500),
                      'NPS_Score': np.random.randint(0,11,500)})
df_nps['NPS_Category'] = pd.cut(df_nps['NPS_Score'], bins=[-1,6,8,10], labels=['Detractor','Passive','Promoter'])
print(df_nps.head())
   CustomerID  Age Region  NPS_Score NPS_Category
0           1   56   West          2    Detractor
1           2   69  North          0    Detractor
2           3   46   East          4    Detractor
3           4   32   West          3    Detractor
4           5   60   East          9     Promoter
ct_nps = pd.crosstab(df_nps['Region'], df_nps['NPS_Category'])
print(ct_nps)
chi2, p, dof, ex = chi2_contingency(ct_nps)
print(f"Chi2: {chi2:.2f}, p-value: {p:.4f}")
NPS_Category  Detractor  Passive  Promoter
Region                                    
East                 76       20        17
North                86       38        25
South                84       19        18
West                 77       15        25
Chi2: 10.14, p-value: 0.1187
sns.countplot(x='Region', hue='NPS_Category', data=df_nps)
plt.title('NPS Category Distribution by Region')
plt.show()
No description has been provided for this image
age_bins = [17,29,39,49,59,70]
labels = ['18-29','30-39','40-49','50-59','60-69']
df_nps['AgeGroup'] = pd.cut(df_nps['Age'], bins=age_bins, labels=labels)
print(df_nps[['Age','AgeGroup']].head())
   Age AgeGroup
0   56    50-59
1   69    60-69
2   46    40-49
3   32    30-39
4   60    60-69
ct_age_nps = pd.crosstab(df_nps['AgeGroup'], df_nps['NPS_Category'])
print(ct_age_nps)
chi2, p, dof, ex = chi2_contingency(ct_age_nps)
print(f"Chi2: {chi2:.2f}, p-value: {p:.4f}")
NPS_Category  Detractor  Passive  Promoter
AgeGroup                                  
18-29                72       17        18
30-39                52       14        15
40-49                74       20        14
50-59                73       15        20
60-69                52       26        18
Chi2: 9.16, p-value: 0.3287
ct_complex = pd.crosstab([df_nps['Region'], df_nps['AgeGroup']], df_nps['NPS_Category'])
print(ct_complex.head())
NPS_Category     Detractor  Passive  Promoter
Region AgeGroup                              
East   18-29            16        2         3
       30-39            10        2         4
       40-49            19        2         0
       50-59            16        5         4
       60-69            15        9         6
ct_stack = pd.crosstab(df_nps['Region'], df_nps['NPS_Category'], normalize='index')
ct_stack.plot(kind='bar', stacked=True, colormap='viridis')
plt.title('NPS Composition by Region (Proportion)')
plt.ylabel('Proportion')
plt.show()
No description has been provided for this image
df_nps_missing = df_nps.copy()
df_nps_missing.loc[::20, 'NPS_Category'] = np.nan
print(df_nps_missing['NPS_Category'].isnull().sum())
ct_missing = pd.crosstab(df_nps_missing['Region'], df_nps_missing['NPS_Category'])
print(ct_missing)
25
NPS_Category  Detractor  Passive  Promoter
Region                                    
East                 71       20        17
North                83       35        25
South                80       18        16
West                 73       14        23
df_filtered = df_nps_missing.dropna(subset=['NPS_Category'])
ct_filled = pd.crosstab(df_filtered['Region'], df_filtered['NPS_Category'])
chi2, p, dof, ex = chi2_contingency(ct_filled)
print(f"After dropping missing: Chi2={chi2:.2f}, p-value={p:.4f}")
After dropping missing: Chi2=8.50, p-value=0.2034
try:
    pd.crosstab(df_nps['Age'], df_nps['Region'])
    print('This produces a huge crosstab with too many groups for age.')
except Exception as e:
    print(str(e))
This produces a huge crosstab with too many groups for age.
try:
    ct_wrong = pd.crosstab(df_nps['Region'], df_nps['NPS_Score'])
    chi2, p, dof, ex = chi2_contingency(ct_wrong)
    print(f"Chi2: {chi2:.2f}, p-value: {p:.4f}")
except Exception as e:
    print('Error:', e)
Chi2: 28.93, p-value: 0.5215
segments = df_nps.groupby('Region')['NPS_Score'].mean().sort_values()
print(segments)
Region
East     4.504425
South    4.719008
West     5.025641
North    5.214765
Name: NPS_Score, dtype: float64
ct = pd.crosstab(df_nps['Region'], df_nps['NPS_Category'], margins=True)
print(ct)
NPS_Category  Detractor  Passive  Promoter  All
Region                                         
East                 76       20        17  113
North                86       38        25  149
South                84       19        18  121
West                 77       15        25  117
All                 323       92        85  500
nps_summary = df_nps.groupby('Region')['NPS_Category'].value_counts(normalize=True).unstack().fillna(0)
nps_summary['NPS_Index'] = (nps_summary['Promoter'] - nps_summary['Detractor'])*100
print(nps_summary[['Promoter','Detractor','NPS_Index']])
NPS_Category  Promoter  Detractor  NPS_Index
Region                                      
East          0.150442   0.672566 -52.212389
North         0.167785   0.577181 -40.939597
South         0.148760   0.694215 -54.545455
West          0.213675   0.658120 -44.444444
# Trend analysis is not applicable here since survey responses are cross-sectional.
print('This dataset does not have a time column for trend analysis!')
This dataset does not have a time column for trend analysis!
ct = pd.crosstab(df['SeniorCitizen'], df['Churn'])
chi2, p, dof, ex = chi2_contingency(ct)
sns.countplot(x='SeniorCitizen', hue='Churn', data=df)
plt.title('Churn by Senior Citizen Status')
plt.show()
print(f"Chi2: {chi2:.2f}, p-value: {p:.4f}")
if p < 0.05:
    print('Recommendation: Offer special retention programs for senior citizens, as churn is significantly related.')
else:
    print('No specific retention strategy needed for seniors.')
No description has been provided for this image
Chi2: 159.43, p-value: 0.0000
Recommendation: Offer special retention programs for senior citizens, as churn is significantly related.

Where to Go Next?#

  • Practice cross-tabs and chi-square tests on your own survey or customer data.
  • Watch our YouTube playlist for more tutorials on survey analytics and market research insight.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.