Lesson 51 · Market Research Analytics in Python
Turning Market Analysis into Business Insights with Python
In this lesson, we explore how real customer and market data is used to answer business questions. We focus on transforming market analysis into actionable…
- CourseMarket Research Analytics in Python
- Lesson51 of 56
- Video18 min
- FormatJupyter notebook · 19 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbTurning Market Analysis into Business Insights#
- In this lesson, we explore how real customer and market data is used to answer business questions.
- We focus on transforming market analysis into actionable insights.
- Understanding customer satisfaction, campaign response, and purchase behavior helps improve products, services, or marketing.
- You will learn how to extract key findings from datasets and deliver clear recommendations.
- Practice converting data patterns into business actions.
import pandas as pd
import numpy as np
import openml
import warnings
warnings.filterwarnings('ignore')
Understanding Customer and Market Datasets#
- We will use authentic customer satisfaction surveys, marketing campaign responses, and online retail data.
- Each dataset captures real-world details, from demographics to open-ended customer feedback.
- Survey data often comes with ranges, ratings, or missing responses.
- Beginners often forget to check for missing values, encoding mistakes, or biased group comparisons.
# Beginner Example 1: Loading Customer Satisfaction Data
dataset = openml.datasets.get_dataset(42178)
df_satisfaction, _, _, _ = dataset.get_data(dataset_format='dataframe')
print(df_satisfaction.shape)
print(df_satisfaction.head(3))
# Beginner Example 2: Loading Online Retail Transaction Data
url = 'https://archive.ics.uci.edu/ml/machine-learning-databases/00502/online_retail_II.xlsx'
df_retail = pd.read_excel(url, sheet_name='Year 2010-2011')
df_retail['InvoiceDate'] = pd.to_datetime(df_retail['InvoiceDate'])
print(df_retail.shape)
print(df_retail.head(3))
# Beginner Example 3: Creating and Loading NPS Survey Data
np.random.seed(42)
df_nps = pd.DataFrame({
'CustomerID': range(1,501),
'Age': np.random.randint(18,70,500),
'Region': np.random.choice(['North','South','East','West'],500),
'NPS_Score': np.random.randint(0,11,500)
})
print(df_nps.shape)
print(df_nps.head(3))
# Intermediate Example 1: Summarizing Customer Satisfaction by Senior Status
satisfaction_by_senior = df_satisfaction.groupby('SeniorCitizen')['Churn'].value_counts(normalize=True).unstack()
print(satisfaction_by_senior)
# Intermediate Example 2: Calculating Repeat Purchase Rate in Retail Data
repeat_customers = df_retail.groupby('Customer ID').size().reset_index(name='Transactions')
repeat_rate = (repeat_customers['Transactions'] > 1).mean()
print(f"Repeat purchase rate: {repeat_rate:.2%}")
# Intermediate Example 3: Summarizing Marketing Campaign Response by Education Level
dataset = openml.datasets.get_dataset(1461)
df_campaign, _, _, _ = dataset.get_data(dataset_format='dataframe')
df_campaign.columns = ['age','job','marital','education','default','balance','housing','loan','contact','day','month',
'duration','campaign','pdays','previous','poutcome','response']
response_by_education = df_campaign.groupby('education')['response'].value_counts(normalize=True).unstack()
print(response_by_education)
# Intermediate Example 4: Average NPS Score by Region
avg_nps_region = df_nps.groupby('Region')['NPS_Score'].mean()
print(avg_nps_region)
# Intermediate Example 5: Analyzing Transaction Value Over Time
df_retail['TotalValue'] = df_retail['Quantity'] * df_retail['Price']
monthly_sales = df_retail.set_index('InvoiceDate').resample('M')['TotalValue'].sum()
print(monthly_sales.tail(6))
# Advanced Example 1: Identifying At-Risk Customers via Satisfaction and Churn
risk_group = df_satisfaction[(df_satisfaction['Churn']=='Yes') & (df_satisfaction['tenure'] <= 12)]
at_risk_pct = len(risk_group) / len(df_satisfaction)
print(f"Percent of new customers at risk of churn: {at_risk_pct:.2%}")
# Advanced Example 2: Calculating Customer Lifetime Value (CLV) for Top Customers
clv = df_retail.groupby('Customer ID')['TotalValue'].sum().sort_values(ascending=False)
top_clv = clv.head(5)
print(top_clv)
# Advanced Example 3: Trend Analysis of NPS Over Time by Region
df_nps['Month'] = np.random.choice(pd.date_range('2022-01-01', periods=12, freq='M'), size=len(df_nps))
nps_trend = df_nps.groupby(['Region','Month'])['NPS_Score'].mean().unstack('Region')
print(nps_trend.tail(6))
# Error Handling Example: Checking for Missing Survey Responses
missing_counts = df_satisfaction.isnull().sum()
print(missing_counts[missing_counts > 0])
# Error Example: Incorrect Aggregation when Grouping by Categorical Columns
try:
wrong_group = df_nps.groupby('NPS_Score').Region.mean()
except Exception as e:
print('Error:', e)
# Error Example: Misinterpreting NPS ScalesCounting 'Promoters' Correctly
promoters = df_nps['NPS_Score'] >= 9
print(f"Percent promoters: {promoters.mean():.2%}")
Best Practices for Market Research Analytics#
- Segment results by customer type, location, and tenure to find patterns.
- Use cross-tabulation to see how two variables interact (e.g., churn by contract type).
- Construct indexes or summary scores for easier business decisions.
- Visualize time trends to catch seasonal effects or sudden changes.
# Pattern Example: Cross-Tabulation of Customer Churn by Contract Type
churn_by_contract = pd.crosstab(df_satisfaction['Contract'], df_satisfaction['Churn'], normalize='index')
print(churn_by_contract)
# Pattern Example: Trend AnalysisMonthly New Customers in Online Retail
df_retail['SignupMonth'] = df_retail['InvoiceDate'].dt.to_period('M')
monthly_new = df_retail.groupby('SignupMonth')['Customer ID'].nunique()
print(monthly_new.tail(6))
Tiny End-to-End Problem: From Survey to Strategic Recommendation#
- Suppose NPS dropped in the West region for two consecutive months.
- What action should management take?
- Analyze the NPS trend, identify the segment, and prepare a short recommendation.
# Solution: Analyze NPS Drop for West Region
nps_trend_west = nps_trend['West']
nps_trend_west_recent = nps_trend_west.tail(3)
print("Recent West Region NPS:")
print(nps_trend_west_recent)
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



