Lesson 58 · Market Research Analytics in Python
Market Insights and Recommendations Case Study with Python Analytics
In this lesson, we will analyze real-world customer and market research data to uncover actionable insights. Understanding market research helps businesses…
- CourseMarket Research Analytics in Python
- Lesson58 of 56
- Video23 min
- FormatJupyter notebook · 22 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbMarket Insights and Recommendations Case Study#
- In this lesson, we will analyze real-world customer and market research data to uncover actionable insights.
- Understanding market research helps businesses make informed decisions that can improve customer satisfaction, retention, and revenue.
- We will practice transforming survey and feedback data into recommendations that support your organization.
import pandas as pd
import numpy as np
import openml
import warnings
warnings.filterwarnings('ignore')
Core Market Research Concepts for This Case Study#
- Market research data often comes from surveys, transactional records, and collected customer feedback.
- Survey data includes demographics, satisfaction ratings, and behavioral indicators.
- Free-text responses provide qualitative insight that numbers alone may not reveal.
- Errors sometimes occur if responses are regrouped wrongly, or if scales like NPS are misinterpreted.
- Understanding what your data represents is essential before giving recommendations.
# Load a customer satisfaction survey dataset
dataset = openml.datasets.get_dataset(42178)
df_cs, _, _, _ = dataset.get_data(dataset_format='dataframe')
print(df_cs.shape)
print(df_cs.head(3))
# Load a marketing campaign performance dataset
dataset = openml.datasets.get_dataset(1461)
df_mc, _, _, _ = dataset.get_data(dataset_format='dataframe')
df_mc.columns = ['age','job','marital','education','default','balance','housing','loan','contact','day','month','duration','campaign','pdays','previous','poutcome','response']
print(df_mc.shape)
print(df_mc.head(3))
# Load a synthetic Net Promoter Score (NPS) survey dataset
np.random.seed(42)
df_nps = pd.DataFrame({
'CustomerID': range(1,501),
'Age': np.random.randint(18,70,500),
'Region': np.random.choice(['North','South','East','West'],500),
'NPS_Score': np.random.randint(0,11,500)
})
print(df_nps.shape)
print(df_nps.head(3))
# Basic aggregation: Average satisfaction by gender
avg_satisfaction = df_cs.groupby('gender')['MonthlyCharges'].mean()
print(avg_satisfaction)
# Cross-tabulation: Churn rate by contract type
ct = pd.crosstab(df_cs['Contract'], df_cs['Churn'], normalize='index')
print(ct)
# Simple NPS calculation: Promoters, Passives, Detractors
promoters = (df_nps['NPS_Score'] >= 9).sum()
detractors = (df_nps['NPS_Score'] <= 6).sum()
passives = ((df_nps['NPS_Score'] > 6) & (df_nps['NPS_Score'] < 9)).sum()
total = len(df_nps)
nps = ((promoters - detractors) / total) * 100
print(f'Net Promoter Score: {nps:.1f}')
# Visualize customer churn by age group
age_bins = [18, 30, 40, 50, 60, 80]
df_cs['age_group'] = pd.cut(df_cs['SeniorCitizen']*30 + 25, bins=age_bins, right=False)
churn_by_age = pd.crosstab(df_cs['age_group'], df_cs['Churn'], normalize='index')
print(churn_by_age)
# Load open-ended customer feedback dataset
df_feedback = pd.DataFrame({
'CustomerID':[1,2,3,4,5],
'Feedback':['Great service and friendly staff','Delivery was slow and packaging was poor','Excellent quality, will buy again','Customer support needs improvement','Good value for money']
})
print(df_feedback)
# Simple sentiment analysis using keyword matching
def simple_sentiment(comment):
positive = ['great', 'excellent', 'good', 'friendly']
negative = ['poor', 'slow', 'needs improvement']
text = comment.lower()
if any(word in text for word in positive):
return 'Positive'
elif any(word in text for word in negative):
return 'Negative'
else:
return 'Neutral'
df_feedback['Sentiment'] = df_feedback['Feedback'].apply(simple_sentiment)
print(df_feedback[['Feedback', 'Sentiment']])
# Find most common NPS score by region
most_common_nps = df_nps.groupby('Region')['NPS_Score'].agg(lambda x: x.value_counts().idxmax())
print(most_common_nps)
# Correlation between contract type and monthly charges
avg_charges_by_contract = df_cs.groupby('Contract')['MonthlyCharges'].mean().sort_values()
print(avg_charges_by_contract)
# Tip: High monthly charges may reveal your best upsell opportunities.
# Handle missing survey responses
missing = df_cs.isnull().sum()
print('Missing values per column:')
print(missing)
# Example of misinterpreting NPS scale
wrong_nps_mean = df_nps['NPS_Score'].mean()
print(f'Incorrect NPS if using mean: {wrong_nps_mean:.2f}')
# Aggregation error: double-counting in groupby
try:
double_count = df_nps.groupby(['Region'])['NPS_Score'].sum().sum()
print('Total NPS scores (possible double count):', double_count)
except Exception as e:
print(e)
# Customer segmentation by contract and churn
segment = df_cs.groupby(['Contract', 'Churn']).size().unstack().fillna(0)
print(segment)
# Trend analysis: Monthly new signups (synthetic cohort data)
np.random.seed(0)
dates = pd.date_range('2021-01-01', periods=24, freq='ME')
df_cohort = pd.DataFrame({
'CustomerID': np.random.randint(1000,2000,len(dates)),
'Signup_Month': dates,
'Active_Users': np.random.randint(50,300,len(dates))
})
print(df_cohort.head(5))
signup_trend = df_cohort.groupby(df_cohort['Signup_Month'].dt.to_period('M'))['Active_Users'].sum()
print(signup_trend)
# Construct a Customer Satisfaction Index
df_cs['MonthlyCharges'] = pd.to_numeric(df_cs['MonthlyCharges'], errors='coerce')
df_cs['TotalCharges'] = pd.to_numeric(df_cs['TotalCharges'], errors='coerce')
df_cs['Satisfaction_Index'] = (
df_cs['MonthlyCharges'] /
df_cs['TotalCharges'].replace(0, np.nan)
) * 100
index_mean = df_cs['Satisfaction_Index'].mean(skipna=True)
print(f"Mean Satisfaction Index: {index_mean:.2f}")
# 1. Load market research data: marketing campaign performance
dataset = openml.datasets.get_dataset(1461)
df_mc, _, _, _ = dataset.get_data(dataset_format='dataframe')
df_mc.columns = ['age','job','marital','education','default','balance','housing','loan','contact','day','month','duration','campaign','pdays','previous','poutcome','response']
# 2. Find customer group most likely to respond positively
response_rate = df_mc.groupby('job')['response'].apply(lambda x: (x == 'yes').mean()).sort_values(ascending=False)
print('Response rate by job:')
print(response_rate)
# 3. Write recommendation to file
with open('campaign_recommendation.txt', 'w') as f:
top_job = response_rate.idxmax()
top_rate = response_rate.max()*100
f.write(f'Target job group: {top_job}\nResponse rate: {top_rate:.2f}%\nRecommendation: Focus next campaign on {top_job}s.')
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



