Lesson 6 · Market Research Analytics in Python
Python Basics for Market Research Analysis: Essential Training
This lesson helps you analyze real customer survey and behavior datasets using Python. Market research insights guide product, campaign, and service…
- CourseMarket Research Analytics in Python
- Lesson6 of 56
- Video22 min
- FormatJupyter notebook · 24 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbPython Basics for Market Research Analysis#
- This lesson helps you analyze real customer survey and behavior datasets using Python.
- Market research insights guide product, campaign, and service decisions.
- You will practice summarizing survey answers, segmenting customers, and finding data-driven insights.
- We focus on business value, not just code.
- By the end, you will extract actionable findings from real market data.
import pandas as pd
import numpy as np
import openml
import warnings
warnings.filterwarnings('ignore')
Core concepts in market research data#
- Surveys capture customer opinions, experiences, or intent.
- Each row usually represents one response or customer.
- Columns can be answers, demographic info, or purchase behavior.
- Ratings (e.g. 1-10 scale), choices, and open-ended text are common.
- Beginners often confuse scales (e.g., interpreting high NPS as bad) or mix up who answered questions.
- Market analysis looks for patterns that drive business action.
# Load customer satisfaction dataset from OpenML
dataset = openml.datasets.get_dataset(42178)
df_cs, _, _, _ = dataset.get_data(dataset_format='dataframe')
print(df_cs.shape)
print(df_cs.head(3))
# Summarize gender distribution
print(df_cs['gender'].value_counts())
# Calculate churn rate (percentage who left)
churn_pct = (df_cs['Churn'] == 'Yes').mean() * 100
print('Customer churn rate: {:.2f}%'.format(churn_pct))
# Load marketing campaign dataset from OpenML
dataset = openml.datasets.get_dataset(1461)
df_mkt, _, _, _ = dataset.get_data(dataset_format='dataframe')
df_mkt.columns = ['age','job','marital','education','default','balance','housing','loan','contact','day','month','duration','campaign','pdays','previous','poutcome','response']
print(df_mkt.shape)
print(df_mkt.head(3))
# Calculate basic response rate
resp_rate = (df_mkt['response'] == 'yes').mean() * 100
print('Overall campaign response rate: {:.2f}%'.format(resp_rate))
# Create a synthetic Net Promoter Score (NPS) survey data
np.random.seed(42)
df_nps = pd.DataFrame({'CustomerID': range(1,501), 'Age': np.random.randint(18,70,500), 'Region': np.random.choice(['North','South','East','West'],500), 'NPS_Score': np.random.randint(0,11,500)})
print(df_nps.head(3))
# Average NPS by region
avg_nps_region = df_nps.groupby('Region')['NPS_Score'].mean()
print('Average NPS by region:')
print(avg_nps_region)
# Classify NPS responders
df_nps['NPS_Type'] = pd.cut(df_nps['NPS_Score'], bins=[-1,6,8,10], labels=['Detractor','Passive','Promoter'])
print(df_nps[['CustomerID','Region','NPS_Score','NPS_Type']].head(5))
import matplotlib.pyplot as plt
# Churn rate by contract type
churn_by_contract = df_cs.groupby('Contract')['Churn'].apply(lambda x: (x=='Yes').mean()*100)
churn_by_contract.plot(kind='bar', color='skyblue', figsize=(6,4))
plt.ylabel('Churn Rate (%)')
plt.title('Churn Rate by Contract Type')
plt.xticks(rotation=15)
plt.tight_layout()
plt.show()
# Find top 5 most common customer jobs
print(df_mkt['job'].value_counts().head(5))
# Cross-tabulate response by housing loan
resp_xtab = pd.crosstab(df_mkt['housing'], df_mkt['response'], margins=True)
print(resp_xtab)
# Simulate missing values in Churn field
df_cs_missing = df_cs.copy()
df_cs_missing.loc[0:9, 'Churn'] = np.nan
print(df_cs_missing['Churn'].isnull().sum())
# Calculate churn rate, skipping missing
valid_churn = df_cs_missing['Churn'].dropna()
true_churn_pct = (valid_churn == 'Yes').mean() * 100
print('Churn rate (excluding missing): {:.2f}%'.format(true_churn_pct))
# Calculate NPS index as (promoters - detractors)/total * 100
n_nps = len(df_nps)
n_promoters = (df_nps['NPS_Type'] == 'Promoter').sum()
n_detractors = (df_nps['NPS_Type'] == 'Detractor').sum()
nps_index = (n_promoters - n_detractors) / n_nps * 100
print('NPS Index: {:.2f}'.format(nps_index))
if nps_index < 0: print('Recommendation: Immediate action needed to address negative customer sentiment.')
elif nps_index < 50: print('Recommendation: Focus on converting passives to promoters.')
else: print('Recommendation: Customer advocacy is strong, maintain current strategy.')
# Bin age for segment analysis
df_mkt['age_group'] = pd.cut(df_mkt['age'], bins=[15,25,35,50,80], labels=['16-25','26-35','36-50','51+'])
resp_by_age = df_mkt.groupby('age_group')['response'].value_counts(normalize=True).unstack().fillna(0)*100
print(resp_by_age)
# Text feedback: simulate a blank entry
df_fb = pd.DataFrame({'CustomerID':[1,2,3,4,5], 'Feedback':['Great service',' ','Excellent quality','Support needs improvement','Good value']})
missing_text = df_fb['Feedback'].str.strip() == ''
print('Missing text feedback:', missing_text.sum())
# Example: wrong NPS logic
wrong_promoters = (df_nps['NPS_Score'] <= 6).sum()
print('Wrongly counted promoters:', wrong_promoters)
# Group by typo - should cause error
try:
resp_wrong = df_mkt.groupby('agge')['response'].mean()
except KeyError as e:
print('Grouping error:', e)
# Segmentation: Contract type and churn
seg = df_cs.groupby(['Contract','Churn']).size().unstack(fill_value=0)
print(seg)
# Cross-tab: payment method and churn
pay_xtab = pd.crosstab(df_cs['PaymentMethod'], df_cs['Churn'], normalize='index').round(2)
print(pay_xtab)
# Create synthetic customer cohort data
np.random.seed(0)
dates = pd.date_range('2021-01-01', periods=24, freq='ME')
df_cohort = pd.DataFrame({'CustomerID': np.random.randint(1000,2000,len(dates)), 'Signup_Month': dates, 'Active_Users': np.random.randint(50,300,len(dates))})
# Trend over time
import matplotlib.pyplot as plt
plt.plot(df_cohort['Signup_Month'], df_cohort['Active_Users'], marker='o')
plt.title('Active Users per Cohort Month')
plt.xlabel('Signup Month')
plt.ylabel('Active Users')
plt.grid(True)
plt.tight_layout()
plt.show()
# End-to-end: Find top drivers of churn among senior citizens
senior = df_cs[df_cs['SeniorCitizen']==1]
senior['MonthlyChargesGroup'] = pd.cut(senior['MonthlyCharges'], bins=[0,50,100,150], labels=['Low','Medium','High'])
by_charge = senior.groupby(['MonthlyChargesGroup'])['Churn'].value_counts(normalize=True).unstack().fillna(0)*100
print(by_charge)
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



