Lesson 1 · Market Research Analytics in Python
Introduction to Market Research and Analytics in Python
In this lesson, we will learn how to analyze real market and customer data. We will focus on how to turn customer feedback and survey responses into…
- CourseMarket Research Analytics in Python
- Lesson1 of 56
- Video21 min
- FormatJupyter notebook · 25 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbIntroduction to Market Research and Analytics#
- In this lesson, we will learn how to analyze real market and customer data.
- We will focus on how to turn customer feedback and survey responses into actionable business insights.
- You will explore different datasets, understand customer needs, and learn key metrics like satisfaction, segmentation, and trends.
- By the end, you will be able to identify opportunities and make recommendations based on customer and market analytics.
import pandas as pd
import numpy as np
import openml
import warnings
warnings.filterwarnings('ignore')
What is Market Research? What Data Will We Use?#
- Market research is the study of customers, competitors, and markets to discover opportunities.
- Customer analytics uses real data from surveys, transactions, and feedback to uncover trends and drivers.
- Our data can include:
- Survey ratings on service satisfaction or NPS (Net Promoter Score).
- Customer demographics like age, gender, and region.
- Open-text responses about quality or experiences.
- Purchase transactions over time.
- Common beginner mistakes include misreading codes, ignoring missing data, and grouping responses incorrectly.
# Beginner Example 1: Load a customer satisfaction survey dataset
dataset = openml.datasets.get_dataset(42178)
df_cs, _, _, _ = dataset.get_data(dataset_format='dataframe')
print(df_cs.shape)
print(df_cs.head(3))
# Beginner Example 2: Summary statistics for survey responses
print(df_cs.describe(include='all'))
# Beginner Example 3: Checking for missing values
missing = df_cs.isnull().sum()
print(missing[missing > 0])
# Beginner Example 4: Load a marketing campaign dataset
dataset2 = openml.datasets.get_dataset(1461)
df_marketing, _, _, _ = dataset2.get_data(dataset_format='dataframe')
df_marketing.columns = ['age','job','marital','education','default','balance','housing','loan','contact','day','month','duration','campaign','pdays','previous','poutcome','response']
print(df_marketing.shape)
print(df_marketing.head(3))
# Beginner Example 5: Count successful campaign responses
success_count = (df_marketing['response'] == 'yes').sum()
print('Number of customers who subscribed:', success_count)
# Beginner Example 6: Load a Net Promoter Score survey dataset (synthetic)
np.random.seed(42)
df_nps = pd.DataFrame({'CustomerID': range(1,501), 'Age': np.random.randint(18,70,500), 'Region': np.random.choice(['North','South','East','West'],500), 'NPS_Score': np.random.randint(0,11,500)})
print(df_nps.shape)
print(df_nps.head(3))
# Intermediate Example 1: Calculate average NPS score by region
avg_nps_by_region = df_nps.groupby('Region')['NPS_Score'].mean()
print(avg_nps_by_region)
# Intermediate Example 2: Identify most common customer job types in the marketing campaign
top_jobs = df_marketing['job'].value_counts().head(5)
print(top_jobs)
# Intermediate Example 3: Segment satisfaction by customer contract type
segment_satisfaction = df_cs.groupby('Contract')['Churn'].value_counts(normalize=True).unstack().fillna(0)
print(segment_satisfaction)
# Intermediate Example 4: Open-ended feedback (synthetic small sample)
df_feedback = pd.DataFrame({'CustomerID':[1,2,3,4,5], 'Feedback':['Great service and friendly staff','Delivery was slow and packaging was poor','Excellent quality, will buy again','Customer support needs improvement','Good value for money']})
print(df_feedback.head(3))
# Intermediate Example 5: Quick sentiment keyword search
positive_words = ['great', 'excellent', 'good']
df_feedback['Is_Positive'] = df_feedback['Feedback'].str.lower().apply(lambda x: any(word in x for word in positive_words))
print(df_feedback[['Feedback', 'Is_Positive']])
# Intermediate Example 6: Load and preview online retail behavior
url = 'https://archive.ics.uci.edu/ml/machine-learning-databases/00502/online_retail_II.xlsx'
df_retail = pd.read_excel(url, sheet_name='Year 2010-2011')
df_retail['InvoiceDate'] = pd.to_datetime(df_retail['InvoiceDate'])
print(df_retail.shape)
print(df_retail.head(3))
# Advanced Example 1: Calculate monthly sales and detect trends
df_retail['Month'] = df_retail['InvoiceDate'].dt.to_period('M')
monthly_sales = df_retail.groupby('Month')['Price'].sum()
print(monthly_sales.tail(12))
# Advanced Example 2: Customer segmentation using satisfaction and contract type
df_cs['SeniorCitizen'] = df_cs['SeniorCitizen'].fillna(0)
segments = df_cs.groupby(['SeniorCitizen', 'Contract'])['Churn'].value_counts(normalize=True).unstack().fillna(0)
print(segments)
# Advanced Example 3: Calculate NPS as a business health indicator
promoters = (df_nps['NPS_Score'] >= 9).sum()
detractors = (df_nps['NPS_Score'] <= 6).sum()
total_responses = len(df_nps)
nps_score = ((promoters - detractors)/total_responses) * 100
print('Overall NPS:', round(nps_score,2))
# Error Handling Example 1: What if survey responses are missing?
df_cs_copy = df_cs.copy()
df_cs_copy.loc[0, 'Churn'] = None
missing_churn = df_cs_copy['Churn'].isnull().sum()
print('Missing Churn responses:', missing_churn)
# Error Handling Example 2: Incorrect grouping (wrong aggregation key)
try:
mistake = df_marketing.groupby('Month')['balance'].mean()
except Exception as e:
print('Error:', e)
# Error Handling Example 3: Misinterpreting NPS scales
if df_nps['NPS_Score'].max() > 10 or df_nps['NPS_Score'].min() < 0:
print('Error: NPS Scores should be between 0 and 10.')
else:
print('NPS Score range is valid.')
# Best Practice 1: Customer segmentation by age and satisfaction
bins = [17, 30, 45, 60, 100]
labels = ['18-30','31-45','46-60','61+']
df_nps['AgeGroup'] = pd.cut(df_nps['Age'], bins=bins, labels=labels)
segment_nps = df_nps.groupby('AgeGroup')['NPS_Score'].mean()
print(segment_nps)
# Best Practice 2: Cross-tabulate churn by contract and payment method
churn_crosstab = pd.crosstab(df_cs['Contract'], df_cs['PaymentMethod'], values=df_cs['Churn'] == 'Yes', aggfunc='mean').fillna(0)
print(churn_crosstab)
# Best Practice 3: Constructing an overall satisfaction index
df_cs['satisfaction_index'] = (df_cs['tenure']/df_cs['tenure'].max())*0.5 + (1 - (df_cs['Churn'] == 'Yes').astype(int))*0.5
print(df_cs[['tenure', 'Churn', 'satisfaction_index']].head(5))
# Best Practice 4: Trend analysis in monthly active users (synthetic cohort data)
np.random.seed(0)
dates = pd.date_range('2021-01-01', periods=24, freq='ME')
df_cohort = pd.DataFrame({'CustomerID': np.random.randint(1000,2000,len(dates)), 'Signup_Month': dates, 'Active_Users': np.random.randint(50,300,len(dates))})
print(df_cohort.head())
trend = df_cohort.set_index('Signup_Month')['Active_Users'].rolling(window=6).mean()
print('6-month rolling average of active users:')
print(trend.dropna().tail())
# End-to-End MARKET RESEARCH mini-project
# 1. Load customer survey data
dataset = openml.datasets.get_dataset(42178)
df, _, _, _ = dataset.get_data(dataset_format='dataframe')
# 2. Calculate churn rate by contract type
churn_rate = df.groupby('Contract')['Churn'].apply(lambda x: (x=='Yes').mean())
print('Churn rate by contract type:')
print(churn_rate)
# 3. Recommend the contract for best retention
lowest_churn = churn_rate.idxmin()
print('\nBest contract for customer retention is:', lowest_churn)
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



