Lesson 53 · Market Research Analytics in Python
Data Storytelling for Stakeholders: Enhance Market Research Analytics Skills
In this lesson, we will learn how to analyze market research and customer analytics data to tell compelling stories to stakeholders. Understanding and…
- CourseMarket Research Analytics in Python
- Lesson53 of 56
- Video24 min
- FormatJupyter notebook · 26 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbData Storytelling for Stakeholders#
- In this lesson, we will learn how to analyze market research and customer analytics data to tell compelling stories to stakeholders.
- Understanding and communicating insights from customer and market data is critical for making better business decisions.
- You will gain hands-on experience using real datasets to extract, visualize, and explain key findings to non-technical business audiences.
- We will practice summarizing customer feedback, segmenting audiences, and creating clear recommendations.
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns
import openml
import warnings
warnings.filterwarnings('ignore')
Market Research Data: Structure and Core Concepts#
- Customer analytics and market research datasets usually contain responses from surveys, demographic information, and behavioral indicators.
- Each row is typically a survey response or a customer record.
- Data can include numeric scores (ex: NPS), categorical ratings (ex: 'Yes'/'No'), and open-ended feedback.
- It is important to check for missing data, inconsistent responses, and understand how responses are coded.
- Common mistakes: forgetting to check how missing data is represented, aggregating data incorrectly, or misunderstanding the meaning of scale-based responses.
dataset = openml.datasets.get_dataset(42178)
df, _, _, _ = dataset.get_data(dataset_format='dataframe')
print(df.shape)
print(df.head(3))
n_missing = df.isnull().sum().sum()
print(f"Total missing values in the survey dataset: {n_missing}")
churn_counts = df['Churn'].value_counts()
print("Customer Churn Counts:")
print(churn_counts)
gender_churn = df.groupby('gender')['Churn'].value_counts().unstack()
print(gender_churn)
churn_rate = df['Churn'].value_counts(normalize=True).loc['Yes']
print(f"Overall churn rate: {churn_rate:.2%}")
sns.countplot(data=df, x='Churn', hue='Contract')
plt.title('Churn by Contract Type')
plt.ylabel('Number of Customers')
plt.xlabel('Churn')
plt.show()
monthly_charges = df.groupby('Churn')['MonthlyCharges'].mean()
print("Average monthly charges by churn status:")
print(monthly_charges)
plt.figure(figsize=(8,4))
sns.boxplot(data=df, x='Churn', y='MonthlyCharges')
plt.title('Monthly Charges: Churned vs Not Churned Customers')
plt.show()
internet_churn = df.groupby('InternetService')['Churn'].value_counts(normalize=True).unstack()
internet_churn['Churn Rate (%)'] = internet_churn['Yes'] * 100
print(internet_churn[['Churn Rate (%)']])
survey_summary = df.describe(include='all')
survey_summary.to_csv('survey_summary.csv')
print('Survey summary statistics saved as survey_summary.csv')
campaign_dataset = openml.datasets.get_dataset(1461)
df_campaign, _, _, _ = campaign_dataset.get_data(dataset_format='dataframe')
df_campaign.columns = ['age','job','marital','education','default','balance','housing','loan','contact','day','month','duration','campaign','pdays','previous','poutcome','response']
pivot = pd.crosstab(df_campaign['job'], df_campaign['response'], normalize='index') * 100
pivot = pivot.rename(columns={'yes':'Yes', 'no':'No'}) if 'yes' in pivot.columns else pivot
print(pivot)
sns.heatmap(pivot, annot=True, fmt='.1f', cmap='coolwarm')
plt.title('Campaign Positive Response Rate by Job Segment (%)')
plt.ylabel('Job')
plt.xlabel('Response')
plt.show()
nps_df = pd.DataFrame({'CustomerID': range(1,501),
'Age': np.random.randint(18,70,500),
'Region': np.random.choice(['North','South','East','West'],500),
'NPS_Score': np.random.randint(0,11,500)})
nps_df['NPS_Type'] = np.where(nps_df['NPS_Score'] >= 9, 'Promoter',
np.where(nps_df['NPS_Score'] <=6, 'Detractor', 'Passive'))
nps_summary = nps_df['NPS_Type'].value_counts(normalize=True) * 100
nps_score = nps_summary['Promoter'] - nps_summary['Detractor']
print('NPS Summary (%) by Type:')
print(nps_summary)
print(f'Overall Net Promoter Score (NPS): {nps_score:.1f}')
region_nps = nps_df.groupby('Region')['NPS_Type'].value_counts(normalize=True).unstack().fillna(0) * 100
region_nps['NPS'] = region_nps['Promoter'] - region_nps['Detractor']
print(region_nps[['Promoter','Passive','Detractor','NPS']])
missing_nps = nps_df['NPS_Score'].isnull().sum()
print(f"Number of missing NPS scores: {missing_nps}")
try:
bad_group = nps_df.groupby('Region')['NPS_Scorez'].mean()
except Exception as e:
print('Error grouping by wrong column name:', e)
if set(nps_df['NPS_Score'].unique()) - set(range(0,11)):
print("Warning: Some NPS scores fall outside the 0-10 range!")
else:
print("All NPS scores are within the valid 0-10 range.")
segmentation = df.groupby(['SeniorCitizen', 'Churn']).size().unstack()
print('Customer Segmentation by Senior Citizen Status and Churn:')
print(segmentation)
cross_tab = pd.crosstab(df['gender'], df['Contract'])
print('Cross-tabulation: Gender vs. Contract Type')
print(cross_tab)
df['Price_Category'] = pd.cut(df['MonthlyCharges'], bins=[0,40,70,150], labels=['Low','Mid','High'])
price_churn = df.groupby('Price_Category')['Churn'].value_counts(normalize=True).unstack().fillna(0) * 100
print(price_churn)
trend = df.groupby('tenure')['Churn'].value_counts(normalize=True).unstack().fillna(0)['Yes'] * 100
plt.figure(figsize=(8,4))
plt.plot(trend.index, trend.values)
plt.title('Churn Rate over Customer Tenure')
plt.xlabel('Tenure (months)')
plt.ylabel('Churn Rate (%)')
plt.show()
# 1. Load open-ended customer feedback
df_feedback = pd.DataFrame({'CustomerID':[1,2,3,4,5],
'Feedback':['Great service and friendly staff','Delivery was slow and packaging was poor','Excellent quality, will buy again','Customer support needs improvement','Good value for money']})
print(df_feedback)
# 2. Simple keyword extraction
keywords = pd.Series(' '.join(df_feedback['Feedback']).lower().split()).value_counts().head(5)
print('Most common feedback keywords:')
print(keywords)
# 3. Write an executive summary
summary = 'Key customer themes: Delivery speed and packaging need improvement, while service and quality are praised. Customers value support and good prices.'
with open('customer_feedback_summary.txt', 'w') as f:
f.write(summary)
print('Summary saved as customer_feedback_summary.txt')
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



