Lesson 32 · Market Research Analytics in Python
Hypothesis Testing in Market Research: Essential Methods and Practical Python Analysis
This lesson teaches you how to use real customer datasets and hypothesis tests to make better business decisions. We focus on comparing customer groups,…
- CourseMarket Research Analytics in Python
- Lesson32 of 56
- Video25 min
- FormatJupyter notebook · 26 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbHypothesis Testing in Market Research#
- This lesson teaches you how to use real customer datasets and hypothesis tests to make better business decisions.
- We focus on comparing customer groups, testing marketing campaigns, and understanding key drivers of satisfaction.
- Hypothesis testing helps you turn raw survey or behavior data into actionable business recommendations.
- You will analyze satisfaction surveys and campaign results step-by-step, using Python and real-world datasets.
import pandas as pd
import numpy as np
import openml
import seaborn as sns
import matplotlib.pyplot as plt
import scipy.stats as stats
import warnings
warnings.filterwarnings('ignore')
Market Research Hypothesis Testing: What, Why, and How#
- In market research, a hypothesis is a business guess about differences or patterns among customers.
- Survey data includes demographic facts, satisfaction ratings, and customer choices.
- Each survey row is a customer and columns describe who they are and what they think or do.
- Beginners often misuse Likert scales as numbers or forget to check group sample sizes.
- Always check for missing values, data type mismatches, and survey weighting before starting your analysis.
# Load the Customer Satisfaction Survey dataset
dataset = openml.datasets.get_dataset(42178)
df_cs, _, _, _ = dataset.get_data(dataset_format='dataframe')
print(df_cs.shape)
print(df_cs.head(3))
# Check for missing survey responses
missing_counts = df_cs.isnull().sum()
print('Missing values per column:')
print(missing_counts[missing_counts > 0])
# What does 'Churn' look like among customers?
churn_counts = df_cs['Churn'].value_counts()
print('Customer churn breakdown:')
print(churn_counts)
# Visualize Churn by Customer Gender
sns.countplot(data=df_cs, x='gender', hue='Churn')
plt.title('Churn by Gender')
plt.xlabel('Gender')
plt.ylabel('Number of Customers')
plt.show()
# Beginner Hypothesis Test: Is churn rate different by gender?
contingency_table = pd.crosstab(df_cs['gender'], df_cs['Churn'])
chi2, p, dof, expected = stats.chi2_contingency(contingency_table)
print('Chi-Square statistic:', chi2)
print('p-value:', p)
# Beginner: Test mean MonthlyCharges by churn group
monthly_charges_churn = df_cs[df_cs['Churn']=='Yes']['MonthlyCharges']
monthly_charges_no = df_cs[df_cs['Churn']=='No']['MonthlyCharges']
t_stat, p_val = stats.ttest_ind(monthly_charges_churn, monthly_charges_no, nan_policy='omit')
print('T-statistic:', t_stat)
print('p-value:', p_val)
print('Average charge (Churned):', monthly_charges_churn.mean())
print('Average charge (Retained):', monthly_charges_no.mean())
# Load marketing campaign dataset
dataset = openml.datasets.get_dataset(1461)
df_mkt, _, _, _ = dataset.get_data(dataset_format='dataframe')
df_mkt.columns = ['age','job','marital','education','default','balance','housing','loan','contact','day','month','duration','campaign','pdays','previous','poutcome','response']
print(df_mkt.shape)
print(df_mkt[['response','job','age']].head(3))
# Check conversion rates by education level
conversion_by_edu = pd.crosstab(df_mkt['education'], df_mkt['response'], normalize='index')
print('Conversion rates by education:')
print(conversion_by_edu)
# Hypothesis test: Is response rate different for 'primary' vs 'tertiary' education?
primary = df_mkt[df_mkt['education'] == 'primary']['response'].apply(lambda x: 1 if x=='yes' else 0)
tertiary = df_mkt[df_mkt['education'] == 'tertiary']['response'].apply(lambda x: 1 if x=='yes' else 0)
t_stat, p_val = stats.ttest_ind(primary, tertiary)
print('T-statistic:', t_stat)
print('p-value:', p_val)
# Visualize age distribution among responders and non-responders
sns.histplot(df_mkt[df_mkt['response']=='yes']['age'], kde=True, color='green', label='Responded', bins=20)
sns.histplot(df_mkt[df_mkt['response']=='no']['age'], kde=True, color='red', label='Not Responded', bins=20)
plt.legend()
plt.title('Age Distribution by Campaign Response')
plt.xlabel('Age')
plt.ylabel('Number of Customers')
plt.show()
# Segment by age group and test if campaign response differs
df_mkt['age_group'] = pd.cut(df_mkt['age'], bins=[18, 30, 45, 65, 100], labels=['18-30','31-45','46-65','66+'])
res_age_table = pd.crosstab(df_mkt['age_group'], df_mkt['response'])
chi2, p, dof, expected = stats.chi2_contingency(res_age_table)
print('Chi-square statistic:', chi2)
print('p-value:', p)
print(res_age_table)
# Advanced: Load Net Promoter Score (NPS) synthetic survey
np.random.seed(42)
df_nps = pd.DataFrame({'CustomerID': range(1,501), 'Age': np.random.randint(18,70,500), 'Region': np.random.choice(['North','South','East','West'],500), 'NPS_Score': np.random.randint(0,11,500)})
print(df_nps.head(3))
# Advanced: Hypothesis test, is mean NPS different by region?
mean_by_region = df_nps.groupby('Region')['NPS_Score'].mean()
f_stat, p_val = stats.f_oneway(*[df_nps[df_nps['Region']==region]['NPS_Score'] for region in df_nps['Region'].unique()])
print('Average NPS by region:')
print(mean_by_region)
print('One-way ANOVA F-statistic:', f_stat)
print('p-value:', p_val)
# Advanced: Proportion of promoters vs detractors by age group
df_nps['category'] = pd.cut(df_nps['NPS_Score'], bins=[-1,6,8,10], labels=['Detractor','Passive','Promoter'])
df_nps['age_group'] = pd.cut(df_nps['Age'], bins=[17,30,45,70], labels=['18-30','31-45','46-70'])
table = pd.crosstab(df_nps['age_group'], df_nps['category'], normalize='index')
print('NPS categories by age group:')
print(table)
# Advanced: Compare NPS category distribution by region
nps_cat_table = pd.crosstab(df_nps['Region'], df_nps['category'])
chi2, p, dof, expected = stats.chi2_contingency(nps_cat_table)
print('Contingency table:')
print(nps_cat_table)
print('Chi-square statistic:', chi2)
print('p-value:', p)
# Error handling: Remove rows with missing NPS scores
n_missing = df_nps['NPS_Score'].isnull().sum()
df_nps_clean = df_nps.dropna(subset=['NPS_Score'])
print(f'Removed {n_missing} rows with missing NPS scores. Remaining: {df_nps_clean.shape[0]}')
# Debugging: Check group sizes before running hypothesis tests
group_sizes = df_nps_clean.groupby('Region').size()
print('Survey responses per region:')
print(group_sizes)
# Debugging: What happens if you misinterpret Likert as numeric?
likert = ['Strongly disagree','Disagree','Neutral','Agree','Strongly agree']
score_map = {'Strongly disagree':1, 'Disagree':2, 'Neutral':3, 'Agree':4, 'Strongly agree':5}
fake_likert = pd.Series(np.random.choice(likert, 30))
numeric_likert = fake_likert.map(score_map)
print('Fake Likert ratings:')
print(fake_likert.head())
print('Numeric encoding:')
print(numeric_likert.head())
print('Mean:', numeric_likert.mean())
# Segmentation: Create customer groups for deeper insights
df_cs['tenure_group'] = pd.cut(df_cs['tenure'], bins=[0,12,24,48,100], labels=['<1yr','1-2yr','2-4yr','4+yr'])
print(df_cs[['tenure','tenure_group']].head(6))
# Cross-tabulation: Churn by payment method
ct = pd.crosstab(df_cs['PaymentMethod'], df_cs['Churn'], normalize='index')
print('Churn rate per payment method:')
print(ct.head())
# Index and score construction: Build a composite satisfaction score
# (Example using available columns; adapt for your data situation)
satisfaction_cols = ['OnlineSecurity','TechSupport','StreamingTV']
df_cs['satisfaction_score'] = (df_cs[satisfaction_cols] == 'Yes').sum(axis=1)
print(df_cs[['satisfaction_score']].describe())
# Trend analysis: Did average monthly charges increase over tenure?
monthly_by_tenure = df_cs.groupby('tenure_group')['MonthlyCharges'].mean()
monthly_by_tenure.plot(kind='bar', color='dodgerblue')
plt.title('Average Monthly Charges by Tenure Group')
plt.xlabel('Tenure Group')
plt.ylabel('Average Monthly Charge')
plt.show()
# End-to-end: Test if satisfaction differs by churn outcome
df_cs['satisfaction_score'] = (df_cs[satisfaction_cols] == 'Yes').sum(axis=1)
churned = df_cs[df_cs['Churn']=='Yes']['satisfaction_score']
retained = df_cs[df_cs['Churn']=='No']['satisfaction_score']
t_stat, p_value = stats.ttest_ind(churned, retained, nan_policy='omit')
print('Satisfaction (churned):', churned.mean())
print('Satisfaction (retained):', retained.mean())
print('t-statistic:', t_stat)
print('p-value:', p_value)
if p_value < 0.05:
print('Conclusion: Satisfaction is significantly different for churned versus retained customers.')
else:
print('Conclusion: Satisfaction difference by churn is not statistically significant.')
# Watch a YouTube guide for more hypothesis test examples
print('For more hands-on examples, search YouTube for "hypothesis testing in customer analytics".')
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



