Lesson 34 · Market Research Analytics in Python
Chi-Square Tests for Analyzing Survey Relationships Using Python | Market Research Analytics
We will learn to analyze whether customer survey responses are related to key demographics or behaviors. This analysis helps businesses discover patterns…
- CourseMarket Research Analytics in Python
- Lesson34 of 56
- Video25 min
- FormatJupyter notebook · 25 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbChi-Square Tests for Survey Relationships#
- We will learn to analyze whether customer survey responses are related to key demographics or behaviors.
- This analysis helps businesses discover patterns such as which customer groups are more likely to churn or recommend a service.
- By applying chi-square tests, we will detect whether variables like satisfaction, churn, or NPS differ across segments.
- The goal is to move from raw survey data to actionable insights for business improvements.
import warnings
warnings.filterwarnings('ignore')
import pandas as pd
import numpy as np
import openml
from scipy.stats import chi2_contingency
import matplotlib.pyplot as plt
import seaborn as sns
Core Concepts: What Are Chi-Square Tests in Market Research?#
- Customer or market surveys collect responses (like satisfaction or NPS) along with demographic or behavioral data.
- Chi-square tests help us check if two categorical variables (such as gender and churn) are statistically related.
- Survey data may be structured with columns like gender, age group, product chosen, and survey outcome.
- Common mistakes include using chi-square on continuous data or not checking group sample sizes.
- Beginner analysts sometimes ignore data cleaning or the risk of empty cells in the data.
dataset = openml.datasets.get_dataset(42178)
df, _, _, _ = dataset.get_data(dataset_format='dataframe')
print(df.shape)
print(df[['gender', 'SeniorCitizen', 'Churn']].head())
ct = pd.crosstab(df['gender'], df['Churn'])
print(ct)
chi2, p, dof, ex = chi2_contingency(ct)
print(f"Chi2 Statistic: {chi2:.3f}, p-value: {p:.4f}")
sns.countplot(x='gender', hue='Churn', data=df)
plt.title('Churn by Gender')
plt.show()
dataset = openml.datasets.get_dataset(1461)
df_marketing, _, _, _ = dataset.get_data(dataset_format='dataframe')
df_marketing.columns = ['age','job','marital','education','default','balance','housing','loan','contact','day','month','duration','campaign','pdays','previous','poutcome','response']
print(df_marketing[['job','response']].head())
ct_job = pd.crosstab(df_marketing['job'], df_marketing['response'])
print(ct_job.head())
chi2, p, dof, ex = chi2_contingency(ct_job)
print(f"Chi2 Statistic: {chi2:.2f}, p-value: {p:.5f}")
if p < 0.05:
print('There is a statistically significant relationship between job type and campaign response.')
else:
print('No significant relationship between job type and campaign response.')
ct_edu = pd.crosstab(df_marketing['education'], df_marketing['response'])
chi2, p, dof, ex = chi2_contingency(ct_edu)
print(ct_edu)
print(f"Chi2: {chi2:.2f}, p-value: {p:.4f}")
np.random.seed(42)
df_nps = pd.DataFrame({'CustomerID': range(1,501),
'Age': np.random.randint(18,70,500),
'Region': np.random.choice(['North','South','East','West'],500),
'NPS_Score': np.random.randint(0,11,500)})
df_nps['NPS_Category'] = pd.cut(df_nps['NPS_Score'], bins=[-1,6,8,10], labels=['Detractor','Passive','Promoter'])
print(df_nps.head())
ct_nps = pd.crosstab(df_nps['Region'], df_nps['NPS_Category'])
print(ct_nps)
chi2, p, dof, ex = chi2_contingency(ct_nps)
print(f"Chi2: {chi2:.2f}, p-value: {p:.4f}")
sns.countplot(x='Region', hue='NPS_Category', data=df_nps)
plt.title('NPS Category Distribution by Region')
plt.show()
age_bins = [17,29,39,49,59,70]
labels = ['18-29','30-39','40-49','50-59','60-69']
df_nps['AgeGroup'] = pd.cut(df_nps['Age'], bins=age_bins, labels=labels)
print(df_nps[['Age','AgeGroup']].head())
ct_age_nps = pd.crosstab(df_nps['AgeGroup'], df_nps['NPS_Category'])
print(ct_age_nps)
chi2, p, dof, ex = chi2_contingency(ct_age_nps)
print(f"Chi2: {chi2:.2f}, p-value: {p:.4f}")
ct_complex = pd.crosstab([df_nps['Region'], df_nps['AgeGroup']], df_nps['NPS_Category'])
print(ct_complex.head())
ct_stack = pd.crosstab(df_nps['Region'], df_nps['NPS_Category'], normalize='index')
ct_stack.plot(kind='bar', stacked=True, colormap='viridis')
plt.title('NPS Composition by Region (Proportion)')
plt.ylabel('Proportion')
plt.show()
df_nps_missing = df_nps.copy()
df_nps_missing.loc[::20, 'NPS_Category'] = np.nan
print(df_nps_missing['NPS_Category'].isnull().sum())
ct_missing = pd.crosstab(df_nps_missing['Region'], df_nps_missing['NPS_Category'])
print(ct_missing)
df_filtered = df_nps_missing.dropna(subset=['NPS_Category'])
ct_filled = pd.crosstab(df_filtered['Region'], df_filtered['NPS_Category'])
chi2, p, dof, ex = chi2_contingency(ct_filled)
print(f"After dropping missing: Chi2={chi2:.2f}, p-value={p:.4f}")
try:
pd.crosstab(df_nps['Age'], df_nps['Region'])
print('This produces a huge crosstab with too many groups for age.')
except Exception as e:
print(str(e))
try:
ct_wrong = pd.crosstab(df_nps['Region'], df_nps['NPS_Score'])
chi2, p, dof, ex = chi2_contingency(ct_wrong)
print(f"Chi2: {chi2:.2f}, p-value: {p:.4f}")
except Exception as e:
print('Error:', e)
segments = df_nps.groupby('Region')['NPS_Score'].mean().sort_values()
print(segments)
ct = pd.crosstab(df_nps['Region'], df_nps['NPS_Category'], margins=True)
print(ct)
nps_summary = df_nps.groupby('Region')['NPS_Category'].value_counts(normalize=True).unstack().fillna(0)
nps_summary['NPS_Index'] = (nps_summary['Promoter'] - nps_summary['Detractor'])*100
print(nps_summary[['Promoter','Detractor','NPS_Index']])
# Trend analysis is not applicable here since survey responses are cross-sectional.
print('This dataset does not have a time column for trend analysis!')
ct = pd.crosstab(df['SeniorCitizen'], df['Churn'])
chi2, p, dof, ex = chi2_contingency(ct)
sns.countplot(x='SeniorCitizen', hue='Churn', data=df)
plt.title('Churn by Senior Citizen Status')
plt.show()
print(f"Chi2: {chi2:.2f}, p-value: {p:.4f}")
if p < 0.05:
print('Recommendation: Offer special retention programs for senior citizens, as churn is significantly related.')
else:
print('No specific retention strategy needed for seniors.')
Where to Go Next?#
- Practice cross-tabs and chi-square tests on your own survey or customer data.
- Watch our YouTube playlist for more tutorials on survey analytics and market research insight.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



