Lesson 4 · Market Research Analytics in Python
Market Research Workflow Using Python: Step-by-Step Training Guide
In this lesson, we will solve real-world market research and customer analytics problems using Python. We will work with genuine business datasets such as…
- CourseMarket Research Analytics in Python
- Lesson4 of 56
- Video22 min
- FormatJupyter notebook · 23 code cells
What you'll learn
- Core Market Research Concepts
- Beginner Example 1: Count Churned vs. Retained Customers
- Beginner Example 2: Basic Summary Statistics for NPS
- Beginner Example 3: Frequency Table for Internet Service Types
- Intermediate Example 1: NPS by Region Segmentation
- Intermediate Example 2: Monthly Customer Retention Trend
- Intermediate Example 3: Campaign Response Analysis
- Advanced Example 1: Constructing NPS Categories
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbMarket Research Workflow Using Python#
- In this lesson, we will solve real-world market research and customer analytics problems using Python.
- We will work with genuine business datasets such as customer satisfaction surveys, campaign data, retail transactions, NPS responses, and customer feedback.
- Understanding and analyzing customer data helps businesses uncover insights, retain customers, and design more effective products and services.
- By the end, you will be able to import survey data, perform key analyses, spot mistakes, and deliver insights, metrics, or recommendations that matter.
import pandas as pd
import numpy as np
import openml
import warnings
warnings.filterwarnings('ignore')
Core Market Research Concepts#
- Market research datasets often contain structured responses: multiple-choice, numeric ratings, text comments, demographic info, purchase history, and more.
- Customer satisfaction, NPS, and campaign datasets capture how different segments respond to products or services.
- Beginners often forget to check for missing or inconsistent responses, or misinterpret scales (e.g., thinking a high Likert score is always good).
- Text feedback is common but must be analyzed differently than numeric data.
- Segmentation and trend analysis are crucial for turning data into business decisions.
dataset = openml.datasets.get_dataset(42178)
df_cs, _, _, _ = dataset.get_data(dataset_format='dataframe')
print(df_cs.shape)
print(df_cs.head(3))
dataset = openml.datasets.get_dataset(1461)
df_mkt, _, _, _ = dataset.get_data(dataset_format='dataframe')
df_mkt.columns = ['age','job','marital','education','default','balance','housing','loan','contact','day','month','duration','campaign','pdays','previous','poutcome','response']
print(df_mkt.shape)
print(df_mkt.head(3))
url = 'https://archive.ics.uci.edu/ml/machine-learning-databases/00502/online_retail_II.xlsx'
df_retail = pd.read_excel(url, sheet_name='Year 2010-2011')
df_retail['InvoiceDate'] = pd.to_datetime(df_retail['InvoiceDate'])
print(df_retail.shape)
print(df_retail.head(3))
np.random.seed(42)
df_nps = pd.DataFrame({'CustomerID': range(1,501), 'Age': np.random.randint(18,70,500), 'Region': np.random.choice(['North','South','East','West'],500), 'NPS_Score': np.random.randint(0,11,500)})
print(df_nps.shape)
print(df_nps.head(3))
df_feedback = pd.DataFrame({'CustomerID':[1,2,3,4,5], 'Feedback':['Great service and friendly staff','Delivery was slow and packaging was poor','Excellent quality, will buy again','Customer support needs improvement','Good value for money']})
print(df_feedback.shape)
print(df_feedback.head(3))
np.random.seed(0)
dates = pd.date_range('2021-01-01', periods=24, freq='ME')
df_cohort = pd.DataFrame({'CustomerID': np.random.randint(1000,2000,len(dates)), 'Signup_Month': dates, 'Active_Users': np.random.randint(50,300,len(dates))})
print(df_cohort.shape)
print(df_cohort.head(3))
Beginner Example 1: Count Churned vs. Retained Customers#
- Identifying churned customers is vital for retention strategy.
- We use the customer satisfaction dataset to count how many left vs. stayed.
churn_counts = df_cs['Churn'].value_counts()
print('Churn breakdown:')
print(churn_counts)
Beginner Example 2: Basic Summary Statistics for NPS#
- NPS scores range from 0 (least likely to recommend) to 10 (most likely).
- High average NPS indicates stronger customer advocacy.
nps_mean = df_nps['NPS_Score'].mean()
nps_std = df_nps['NPS_Score'].std()
print(f'Average NPS: {nps_mean:.2f}')
print(f'Standard deviation of NPS: {nps_std:.2f}')
Beginner Example 3: Frequency Table for Internet Service Types#
- Segmenting service types reveals what options customers choose most.
- This helps design future offers and communications.
value_freq = df_cs['InternetService'].value_counts()
print('Frequency of Internet Service Types:')
print(value_freq)
Intermediate Example 1: NPS by Region Segmentation#
- Breaking down NPS by geography spots regional strengths and weaknesses.
- You can target support or marketing by location.
nps_region = df_nps.groupby('Region')['NPS_Score'].mean()
print('Average NPS by Region:')
print(nps_region)
Intermediate Example 2: Monthly Customer Retention Trend#
- Tracking active users over time shows retention performance.
- It helps reveal effects of product changes or campaigns.
trend = df_cohort.set_index('Signup_Month')['Active_Users']
print('Monthly Active Users:')
print(trend)
Intermediate Example 3: Campaign Response Analysis#
- Understanding who responds to marketing campaigns increases ROI.
- We measure mean account balance by response in the campaign data.
mean_balance = df_mkt.groupby('response')['balance'].mean()
print('Average Balance by Campaign Response:')
print(mean_balance)
Advanced Example 1: Constructing NPS Categories#
- You must segment NPS: 0-6 = Detractor, 7-8 = Passive, 9-10 = Promoter.
- This is the industry standard for NPS reporting.
bins = [0,6,8,10]
labels = ['Detractor','Passive','Promoter']
df_nps['NPS_Category'] = pd.cut(df_nps['NPS_Score'], bins=[-1,6,8,10], labels=labels)
cat_counts = df_nps['NPS_Category'].value_counts()
print('Counts by NPS Category:')
print(cat_counts)
Advanced Example 2: Open-Ended Feedback Sentiment Keyword Search#
- Customers describe good and bad experiences in their own words.
- Finding keywords related to issues or praise helps prioritize business actions.
keyword = 'improvement'
hits = df_feedback['Feedback'].str.lower().str.contains(keyword)
print(f'Customers mentioning "{keyword}":')
print(df_feedback[hits])
missing = df_cs.isnull().sum()
print('Missing survey responses per column:')
print(missing[missing > 0])
grouped = df_mkt.groupby('job')['balance'].sum()
print('Total balance by job:')
print(grouped)
incorrect_total = grouped.sum()
dataset_total = df_mkt['balance'].sum()
print(f'Check: Grouped total = {incorrect_total}, Actual total = {dataset_total}')
nps_likert = [0,1,2,3,4,5,6,7,8,9,10]
for val in nps_likert:
if val <= 6:
category = 'Detractor'
elif val <= 8:
category = 'Passive'
else:
category = 'Promoter'
print(f'NPS {val}: {category}')
Best Practices: Segment, Cross-Tab, and Track Trends#
- Always check for missing data and results that make business sense.
- Segment by region, age, or service type to uncover hidden opportunities.
- Cross-tabulate customer demographics with behavior for deeper insights.
- Build simple indexes like Customer Satisfaction or NPS consistently.
- Track key metrics (like retention or NPS) monthly to diagnose change.
ctab = pd.crosstab(df_cs['SeniorCitizen'], df_cs['InternetService'])
print('Cross-tabulation: Senior Citizen vs. Internet Service')
print(ctab)
segmentation = df_nps.groupby(['Region', 'NPS_Category']).size().unstack(fill_value=0)
print('NPS Segmentation by Region:')
print(segmentation)
monthly_charges_trend = df_cs.groupby('tenure')['MonthlyCharges'].mean()
print('Average Monthly Charges by Tenure:')
print(monthly_charges_trend.head())
End-to-End Mini Project: Identify At-Risk Customer Segment#
- Goal: Use survey and behavior data to spot a customer group with high churn risk.
- Step 1: Find all senior citizens with fiber optic internet who have churned.
- Step 2: Count and recommend a next step for the business.
- This is a classic real-world task every market research analyst performs.
at_risk = df_cs[(df_cs['SeniorCitizen'] == 1) & (df_cs['InternetService'] == 'Fiber optic') & (df_cs['Churn'] == 'Yes')]
count_risk = at_risk.shape[0]
print(f'Number of at-risk senior fiber optic customers who churned: {count_risk}')
if count_risk > 0:
print('Action Recommendation: Launch a retention campaign or survey for these customers.')
else:
print('No at-risk customers found in this segment.')
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



