Lesson 2 · Market Research Analytics in Python
Primary vs Secondary Market Research Data Explained for Analytics Training
In this lesson, we will explore how to analyze primary and secondary market research data using Python. Understanding these data types helps businesses make…
- CourseMarket Research Analytics in Python
- Lesson2 of 56
- Video20 min
- FormatJupyter notebook · 17 code cells
What you'll learn
- Core Market Research Data Concepts
- Beginner Example 1: Summarizing Survey Scores
- Beginner Example 2: Proportion of Satisfied Customers
- Beginner Example 3: Survey Responses by Region
- Intermediate Example 1: NPS Distribution Histogram
- Intermediate Example 2: Segment Churn Rate by Contract Type
- Intermediate Example 3: Calculating Customer Lifetime Value Features
- Advanced Example 1: Identify Customer Segments from NPS
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbPrimary vs Secondary Market Research Data#
- In this lesson, we will explore how to analyze primary and secondary market research data using Python.
- Understanding these data types helps businesses make informed decisions based on survey results and existing records.
- You will learn how to recognize, process, and generate customer insight from real survey and behavioral datasets.
- We will produce practical, business-ready analytics from both primary (original) and secondary (existing) market data.
import warnings
warnings.filterwarnings('ignore')
import pandas as pd
import numpy as np
import openml
Core Market Research Data Concepts#
- Primary data comes directly from original sources such as customer surveys or interviews.
- Secondary data is collected for another purpose but can be reused, like company records or public datasets.
- Typical datasets include demographics, customer responses, purchase history, and open text feedback.
- Beginners often confuse primary with secondary data, or treat survey answers as exact truths instead of opinions.
- It is important to check for missing values, poorly designed questions, or misinterpretation of scaled responses.
# Load a primary data example: synthetic Net Promoter Score (NPS) survey
np.random.seed(42)
nps_df = pd.DataFrame({
'CustomerID': range(1, 501),
'Age': np.random.randint(18, 70, 500),
'Region': np.random.choice(['North', 'South', 'East', 'West'], 500),
'NPS_Score': np.random.randint(0, 11, 500)
})
print(nps_df.head(3))
# Load a secondary data example: OpenML customer satisfaction survey data
dataset = openml.datasets.get_dataset(42178)
csat_df, _, _, _ = dataset.get_data(dataset_format='dataframe')
print(csat_df.shape)
print(csat_df.head(3))
# Load a behavioral/secondary data example: online retail transactions from UCI
url = 'https://archive.ics.uci.edu/ml/machine-learning-databases/00502/online_retail_II.xlsx'
retail_df = pd.read_excel(url, sheet_name='Year 2010-2011')
retail_df['InvoiceDate'] = pd.to_datetime(retail_df['InvoiceDate'])
print(retail_df.head(3))
Beginner Example 1: Summarizing Survey Scores#
- Let us start simple: calculate the average Net Promoter Score (NPS) for all respondents.
- NPS is a common primary research metric for customer loyalty.
avg_nps = nps_df['NPS_Score'].mean()
print(f'Overall mean NPS score: {avg_nps:.2f}')
Beginner Example 2: Proportion of Satisfied Customers#
- Using secondary data, let us find out what portion of customers are classified as 'Churn = No'.
- This helps us understand how many customers are staying, using existing survey records.
satisfied_prop = (csat_df['Churn'] == 'No').mean()
print(f'Proportion of customers who did not churn: {satisfied_prop:.2%}')
Beginner Example 3: Survey Responses by Region#
- Let us look at how average NPS scores vary by customer region, a common segmentation in customer analytics.
region_nps = nps_df.groupby('Region')['NPS_Score'].mean()
print(region_nps)
Intermediate Example 1: NPS Distribution Histogram#
- Visualize the spread of NPS scores in our primary survey using a simple plot.
- Spot any bias or gaps in how customers rate their experience.
import matplotlib.pyplot as plt
plt.hist(nps_df['NPS_Score'], bins=11, edgecolor='k')
plt.xlabel('NPS Score')
plt.ylabel('Frequency')
plt.title('Distribution of NPS Survey Scores')
plt.show()
Intermediate Example 2: Segment Churn Rate by Contract Type#
- Use secondary data to find which contract type has the highest customer loss rate.
churn_by_contract = csat_df.groupby('Contract')['Churn'].value_counts(normalize=True).unstack()['Yes']
print(churn_by_contract)
Intermediate Example 3: Calculating Customer Lifetime Value Features#
- Let us use transaction data to estimate the average order size per customer as a business metric.
retail_df_nonnull = retail_df.dropna(subset=['Customer ID'])
order_values = retail_df_nonnull.groupby('Customer ID').apply(lambda g: (g['Quantity'] * g['Price']).sum())
avg_order_value = order_values.mean()
print(f'Average order size per customer: GBP {avg_order_value:.2f}')
Advanced Example 1: Identify Customer Segments from NPS#
- Segments can be based on NPS: Promoters (9-10), Passives (7-8), Detractors (0-6).
- Create a new segment feature and compare their counts.
def nps_segment(score):
if score >= 9: return 'Promoter'
elif score >= 7: return 'Passive'
else: return 'Detractor'
nps_df['Segment'] = nps_df['NPS_Score'].apply(nps_segment)
print(nps_df['Segment'].value_counts())
Advanced Example 2: Cross-tabulation of Churn by Paperless Billing#
- Cross-tabulations reveal relationships between two categorical features, such as billing preference and churn.
crosstab = pd.crosstab(csat_df['PaperlessBilling'], csat_df['Churn'], margins=True, normalize='index')
print(crosstab)
# Error Handling 1: Find Missing Survey Responses
missing_nps = nps_df.isnull().sum()
print('Missing values by column:\n', missing_nps)
# Error Handling 2: Check for Incorrect Aggregation (Grouping by wrong column)
try:
wrong_group = nps_df.groupby('NPS_Score')['Region'].count()
print(wrong_group.head())
except Exception as e:
print(f'Error: {e}')
# Error Handling 3: Misinterpreting Likert Scales or NPS Scores
if nps_df['NPS_Score'].min() < 0 or nps_df['NPS_Score'].max() > 10:
print('Warning: NPS scores must be between 0 and 10 (inclusive)!')
else:
print('NPS score range is valid.')
Market Research Analytics Patterns#
- Segmenting customers helps businesses target specific groups and personalize marketing.
- Cross-tabs are used to explore relationships (e.g., does paperless billing drive retention?).
- Constructing indiceslike average satisfaction or NPSsummarizes broad trends without overcomplicating.
- Trend analysis helps spot changes over time in satisfaction, revenue, or loyalty.
# Pattern: Time Trend Analysis of Order Value (secondary behavioral data)
retail_df_nonnull['Month'] = retail_df_nonnull['InvoiceDate'].dt.to_period('M')
monthly_sales = retail_df_nonnull.groupby('Month').apply(lambda g: (g['Quantity'] * g['Price']).sum())
monthly_sales.plot(kind='line', marker='o')
plt.ylabel('Total Sales (GBP)')
plt.title('Monthly Sales Trend (2010-2011)')
plt.show()
End-to-End Business Example: From Raw Survey to Recommendation#
- Task: Find the region with the most Promoters in our NPS survey, and recommend a regional marketing action.
- Step 1: Calculate number of Promoters per region.
- Step 2: Identify the top region.
- Step 3: Suggest an action for business growth.
promoter_counts = nps_df[nps_df['Segment'] == 'Promoter'].groupby('Region').size()
top_region = promoter_counts.idxmax()
print(f'Region with most Promoters: {top_region}')
print('Recommendation: Focus new referral campaigns in this region to build on current loyalty.')
YouTube: Learn More Market Research Analytics#
- For more hands-on tips, check out YouTubesearch for 'Python market research analytics' and follow along with examples.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



