Lesson 11 · Market Research Analytics in Python
Mastering Survey & Market Research Dataset Structures in Python | Analytics Training
Understand common structures in real-world survey and market research data Learn why dataset structure matters for business decision-making Gain insight…
- CourseMarket Research Analytics in Python
- Lesson11 of 56
- Video19 min
- FormatJupyter notebook · 15 code cells
What you'll learn
- What are Survey and Market Research Datasets?
- Beginner Example 1: Loading a Customer Satisfaction Survey Dataset
- Beginner Example 2: Examining Common Survey Column Types
- Beginner Example 3: Loading a Marketing Campaign Dataset
- Intermediate Example 1: Loading an Online Retail Behavior Dataset
- Intermediate Example 2: Handling Net Promoter Score (NPS) Survey Data
- Intermediate Example 3: Open-Ended Customer Feedback Data
- Advanced Example 1: Understanding Customer Cohort Datasets
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbStructure of Survey and Market Research Datasets#
- Understand common structures in real-world survey and market research data
- Learn why dataset structure matters for business decision-making
- Gain insight into best dataset formats for analytics
- Identify the most common mistakes in organizing and analyzing survey responses
- Discover how to explore, clean, and summarize typical market research datasets
import pandas as pd
import numpy as np
import openml
import warnings
warnings.filterwarnings('ignore')
What are Survey and Market Research Datasets?#
- These datasets are structured tables that collect responses from people, such as customers or survey participants.
- Columns may include demographics, responses to questions, ratings, open-ended feedback, or purchasing behavior.
- Market research datasets often combine business outcomes (like Churn or Purchase) with survey metrics or experimental results.
- Common mistakes include treating categorical answers as numeric, ignoring missing data, or misinterpreting scaled responses.
- Well-structured data helps you find patterns, segment users, and make business recommendations.
Beginner Example 1: Loading a Customer Satisfaction Survey Dataset#
- We start by loading a real customer satisfaction survey.
- Let us preview its structure and discuss its columns.
# Load customer satisfaction survey from OpenML
dataset = openml.datasets.get_dataset(42178)
df_satis, _, _, _ = dataset.get_data(dataset_format='dataframe')
print(df_satis.shape)
print(df_satis.head(3))
Beginner Example 2: Examining Common Survey Column Types#
- Survey datasets often include categorical, numeric, and textual columns.
- Let us look at typical columns from our customer satisfaction data.
- Columns include: Gender, SeniorCitizen, Partner, Dependents, tenure, PhoneService, etc.
print('Columns:')
print(', '.join(df_satis.columns))
print('\nExample data:')
print(df_satis.iloc[0])
Beginner Example 3: Loading a Marketing Campaign Dataset#
- Let us load real customer response data from a bank marketing campaign.
- Notice how it is structured differently than a satisfaction survey.
dataset = openml.datasets.get_dataset(1461)
df_marketing, _, _, _ = dataset.get_data(dataset_format='dataframe')
df_marketing.columns = ['age','job','marital','education','default','balance','housing','loan','contact','day','month','duration','campaign','pdays','previous','poutcome','response']
print(df_marketing.shape)
print(df_marketing.head(3))
Intermediate Example 1: Loading an Online Retail Behavior Dataset#
- Survey data may be combined with purchasing or behavioral records.
- Let us load and preview an online retail transactions dataset.
url = 'https://archive.ics.uci.edu/ml/machine-learning-databases/00502/online_retail_II.xlsx'
df_retail = pd.read_excel(url, sheet_name='Year 2010-2011')
df_retail['InvoiceDate'] = pd.to_datetime(df_retail['InvoiceDate'])
print(df_retail.shape)
print(df_retail.head(3))
Intermediate Example 2: Handling Net Promoter Score (NPS) Survey Data#
- NPS surveys rate customer loyalty from 0 (not likely) to 10 (very likely).
- We create a sample NPS survey dataset and examine its structure.
np.random.seed(42)
df_nps = pd.DataFrame({'CustomerID': range(1,501), 'Age': np.random.randint(18,70,500), 'Region': np.random.choice(['North','South','East','West'],500), 'NPS_Score': np.random.randint(0,11,500)})
print(df_nps.shape)
print(df_nps.head(3))
Intermediate Example 3: Open-Ended Customer Feedback Data#
- Some surveys include free-text responses for richer insights.
- Let us inspect an example open-ended feedback dataset.
df_feedback = pd.DataFrame({'CustomerID':[1,2,3,4,5], 'Feedback':['Great service and friendly staff','Delivery was slow and packaging was poor','Excellent quality, will buy again','Customer support needs improvement','Good value for money']})
print(df_feedback.shape)
print(df_feedback.head(3))
Advanced Example 1: Understanding Customer Cohort Datasets#
- Cohort datasets group customers by signup date or behavior period, tracking retention.
- Let us generate a synthetic cohort table and discuss how it structures time-based analysis.
np.random.seed(0)
dates = pd.date_range('2021-01-01', periods=24, freq='ME')
df_cohort = pd.DataFrame({'CustomerID': np.random.randint(1000,2000,len(dates)), 'Signup_Month': dates, 'Active_Users': np.random.randint(50,300,len(dates))})
print(df_cohort.shape)
print(df_cohort.head(3))
Advanced Example 3: Multi-table Structure and Data Dictionary Creation#
- Complex research projects organize datasets across multiple related tables.
- Creating a data dictionary helps everyone understand variables and formats.
data_dict = pd.DataFrame({'Column': df_satis.columns, 'Type': df_satis.dtypes.astype(str), 'Example_Value': df_satis.iloc[0].astype(str).values})
print(data_dict)
Error Handling Example 1: Detecting Missing Survey Responses#
- Real survey datasets often contain blank or missing answers.
- Let us look for missing data in the satisfaction survey.
missing_counts = df_satis.isnull().sum()
print('Columns with missing values:')
print(missing_counts[missing_counts > 0])
Error Handling Example 2: Incorrect Grouping/Segmentation#
- A common mistake is grouping by the wrong field, leading to bad insights.
- Let us see what happens if we group NPS results by an unrelated column.
try:
print(df_nps.groupby('CustomerID')['NPS_Score'].mean().head())
except Exception as e:
print(f'Error: {e}')
Error Handling Example 3: Misinterpreting Scaled Responses#
- Likert or NPS scales are often numeric but should be analyzed as ordered categories.
- Let us look at the unique survey values and summary statistics.
unique_nps = df_nps['NPS_Score'].unique()
print(sorted(unique_nps))
print(df_nps['NPS_Score'].describe())
Best Practices: Customer Segmentation, Cross Tabulation, and Score Construction#
- Segmenting customers by demographics or behavior unlocks actionable patterns.
- Cross-tabulating responses with features highlights differences between groups.
- Constructing indices or summary scores creates easy-to-read dashboards.
- Let us segment marketing campaign responses by job type and education.
cross_tab = pd.crosstab(df_marketing['job'], df_marketing['education'])
print(cross_tab)
Best Practices: Time-based Trend Analysis with Cohorts#
- Trend analysis shows if metrics improve over time after a product launch or campaign.
- Let us plot active user counts by cohort signup month.
import matplotlib.pyplot as plt
plt.figure(figsize=(10,5))
plt.plot(df_cohort['Signup_Month'], df_cohort['Active_Users'], marker='o')
plt.xlabel('Signup Month')
plt.ylabel('Active Users')
plt.title('Active Users by Cohort Signup Month')
plt.xticks(rotation=45)
plt.tight_layout()
plt.show()
End-to-End Example: From Survey Structure to Business Insight#
- Let us walk through a simple market research workflow using survey data.
- Question: Are senior citizens more likely to churn than younger customers?
- Steps: Segment, summarize churn rates, and make a business recommendation.
grouped = df_satis.groupby('SeniorCitizen')['Churn'].value_counts(normalize=True).unstack().fillna(0)
print(grouped)
if 'Yes' in grouped.columns:
senior_churn = grouped.loc[1, 'Yes']
non_senior_churn = grouped.loc[0, 'Yes']
print(f'Senior citizen churn rate: {senior_churn:.2%}')
print(f'Non-senior citizen churn rate: {non_senior_churn:.2%}')
print('Business Insight: Focus retention efforts on senior citizens if their churn rate is higher.')
Recap: Structure Enables Effective Market Research Analysis#
- Well-structured survey and customer data let you segment, aggregate, and extract actionable business stories.
- Always check for missing data, column types, and correct segmentation.
- Start every project with clear data documentation and a simple analysis outline.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



