Mathew K Analytics

Lesson 2 · Market Research Analytics in Python

Primary vs Secondary Market Research Data Explained for Analytics Training

In this lesson, we will explore how to analyze primary and secondary market research data using Python. Understanding these data types helps businesses make…

⬇ Download notebookOpen in Colab ↗

📓 Full notebook

Download .ipynb

Primary vs Secondary Market Research Data#

  • In this lesson, we will explore how to analyze primary and secondary market research data using Python.
  • Understanding these data types helps businesses make informed decisions based on survey results and existing records.
  • You will learn how to recognize, process, and generate customer insight from real survey and behavioral datasets.
  • We will produce practical, business-ready analytics from both primary (original) and secondary (existing) market data.
import warnings
warnings.filterwarnings('ignore')
import pandas as pd
import numpy as np
import openml

Core Market Research Data Concepts#

  • Primary data comes directly from original sources such as customer surveys or interviews.
  • Secondary data is collected for another purpose but can be reused, like company records or public datasets.
  • Typical datasets include demographics, customer responses, purchase history, and open text feedback.
  • Beginners often confuse primary with secondary data, or treat survey answers as exact truths instead of opinions.
  • It is important to check for missing values, poorly designed questions, or misinterpretation of scaled responses.
# Load a primary data example: synthetic Net Promoter Score (NPS) survey
np.random.seed(42)
nps_df = pd.DataFrame({
    'CustomerID': range(1, 501),
    'Age': np.random.randint(18, 70, 500),
    'Region': np.random.choice(['North', 'South', 'East', 'West'], 500),
    'NPS_Score': np.random.randint(0, 11, 500)
})
print(nps_df.head(3))
   CustomerID  Age Region  NPS_Score
0           1   56   West          2
1           2   69  North          0
2           3   46   East          4
# Load a secondary data example: OpenML customer satisfaction survey data
dataset = openml.datasets.get_dataset(42178)
csat_df, _, _, _ = dataset.get_data(dataset_format='dataframe')
print(csat_df.shape)
print(csat_df.head(3))
(7043, 20)
   gender  SeniorCitizen Partner Dependents  tenure PhoneService  \
0  Female              0     Yes         No       1           No   
1    Male              0      No         No      34          Yes   
2    Male              0      No         No       2          Yes   

      MultipleLines InternetService OnlineSecurity OnlineBackup  \
0  No phone service             DSL             No          Yes   
1                No             DSL            Yes           No   
2                No             DSL            Yes          Yes   

  DeviceProtection TechSupport StreamingTV StreamingMovies        Contract  \
0               No          No          No              No  Month-to-month   
1              Yes          No          No              No        One year   
2               No          No          No              No  Month-to-month   

  PaperlessBilling     PaymentMethod  MonthlyCharges TotalCharges Churn  
0              Yes  Electronic check           29.85        29.85    No  
1               No      Mailed check           56.95       1889.5    No  
2              Yes      Mailed check           53.85       108.15   Yes  
# Load a behavioral/secondary data example: online retail transactions from UCI
url = 'https://archive.ics.uci.edu/ml/machine-learning-databases/00502/online_retail_II.xlsx'
retail_df = pd.read_excel(url, sheet_name='Year 2010-2011')
retail_df['InvoiceDate'] = pd.to_datetime(retail_df['InvoiceDate'])
print(retail_df.head(3))
  Invoice StockCode                         Description  Quantity  \
0  536365    85123A  WHITE HANGING HEART T-LIGHT HOLDER         6   
1  536365     71053                 WHITE METAL LANTERN         6   
2  536365    84406B      CREAM CUPID HEARTS COAT HANGER         8   

          InvoiceDate  Price  Customer ID         Country  
0 2010-12-01 08:26:00   2.55      17850.0  United Kingdom  
1 2010-12-01 08:26:00   3.39      17850.0  United Kingdom  
2 2010-12-01 08:26:00   2.75      17850.0  United Kingdom  

Beginner Example 1: Summarizing Survey Scores#

  • Let us start simple: calculate the average Net Promoter Score (NPS) for all respondents.
  • NPS is a common primary research metric for customer loyalty.
avg_nps = nps_df['NPS_Score'].mean()
print(f'Overall mean NPS score: {avg_nps:.2f}')
Overall mean NPS score: 4.89

Beginner Example 2: Proportion of Satisfied Customers#

  • Using secondary data, let us find out what portion of customers are classified as 'Churn = No'.
  • This helps us understand how many customers are staying, using existing survey records.
satisfied_prop = (csat_df['Churn'] == 'No').mean()
print(f'Proportion of customers who did not churn: {satisfied_prop:.2%}')
Proportion of customers who did not churn: 73.46%

Beginner Example 3: Survey Responses by Region#

  • Let us look at how average NPS scores vary by customer region, a common segmentation in customer analytics.
region_nps = nps_df.groupby('Region')['NPS_Score'].mean()
print(region_nps)
Region
East     4.504425
North    5.214765
South    4.719008
West     5.025641
Name: NPS_Score, dtype: float64

Intermediate Example 1: NPS Distribution Histogram#

  • Visualize the spread of NPS scores in our primary survey using a simple plot.
  • Spot any bias or gaps in how customers rate their experience.
import matplotlib.pyplot as plt
plt.hist(nps_df['NPS_Score'], bins=11, edgecolor='k')
plt.xlabel('NPS Score')
plt.ylabel('Frequency')
plt.title('Distribution of NPS Survey Scores')
plt.show()
No description has been provided for this image

Intermediate Example 2: Segment Churn Rate by Contract Type#

  • Use secondary data to find which contract type has the highest customer loss rate.
churn_by_contract = csat_df.groupby('Contract')['Churn'].value_counts(normalize=True).unstack()['Yes']
print(churn_by_contract)
Contract
Month-to-month    0.427097
One year          0.112695
Two year          0.028319
Name: Yes, dtype: float64

Intermediate Example 3: Calculating Customer Lifetime Value Features#

  • Let us use transaction data to estimate the average order size per customer as a business metric.
retail_df_nonnull = retail_df.dropna(subset=['Customer ID'])
order_values = retail_df_nonnull.groupby('Customer ID').apply(lambda g: (g['Quantity'] * g['Price']).sum())
avg_order_value = order_values.mean()
print(f'Average order size per customer: GBP {avg_order_value:.2f}')
Average order size per customer: GBP 1898.46

Advanced Example 1: Identify Customer Segments from NPS#

  • Segments can be based on NPS: Promoters (9-10), Passives (7-8), Detractors (0-6).
  • Create a new segment feature and compare their counts.
def nps_segment(score):
    if score >= 9: return 'Promoter'
    elif score >= 7: return 'Passive'
    else: return 'Detractor'
nps_df['Segment'] = nps_df['NPS_Score'].apply(nps_segment)
print(nps_df['Segment'].value_counts())
Segment
Detractor    323
Passive       92
Promoter      85
Name: count, dtype: int64

Advanced Example 2: Cross-tabulation of Churn by Paperless Billing#

  • Cross-tabulations reveal relationships between two categorical features, such as billing preference and churn.
crosstab = pd.crosstab(csat_df['PaperlessBilling'], csat_df['Churn'], margins=True, normalize='index')
print(crosstab)
Churn                   No       Yes
PaperlessBilling                    
No                0.836699  0.163301
Yes               0.664349  0.335651
All               0.734630  0.265370
# Error Handling 1: Find Missing Survey Responses
missing_nps = nps_df.isnull().sum()
print('Missing values by column:\n', missing_nps)
Missing values by column:
 CustomerID    0
Age           0
Region        0
NPS_Score     0
Segment       0
dtype: int64
# Error Handling 2: Check for Incorrect Aggregation (Grouping by wrong column)
try:
    wrong_group = nps_df.groupby('NPS_Score')['Region'].count()
    print(wrong_group.head())
except Exception as e:
    print(f'Error: {e}')
NPS_Score
0    52
1    39
2    50
3    45
4    45
Name: Region, dtype: int64
# Error Handling 3: Misinterpreting Likert Scales or NPS Scores
if nps_df['NPS_Score'].min() < 0 or nps_df['NPS_Score'].max() > 10:
    print('Warning: NPS scores must be between 0 and 10 (inclusive)!')
else:
    print('NPS score range is valid.')
NPS score range is valid.

Market Research Analytics Patterns#

  • Segmenting customers helps businesses target specific groups and personalize marketing.
  • Cross-tabs are used to explore relationships (e.g., does paperless billing drive retention?).
  • Constructing indiceslike average satisfaction or NPSsummarizes broad trends without overcomplicating.
  • Trend analysis helps spot changes over time in satisfaction, revenue, or loyalty.
# Pattern: Time Trend Analysis of Order Value (secondary behavioral data)
retail_df_nonnull['Month'] = retail_df_nonnull['InvoiceDate'].dt.to_period('M')
monthly_sales = retail_df_nonnull.groupby('Month').apply(lambda g: (g['Quantity'] * g['Price']).sum())
monthly_sales.plot(kind='line', marker='o')
plt.ylabel('Total Sales (GBP)')
plt.title('Monthly Sales Trend (2010-2011)')
plt.show()
No description has been provided for this image

End-to-End Business Example: From Raw Survey to Recommendation#

  • Task: Find the region with the most Promoters in our NPS survey, and recommend a regional marketing action.
  • Step 1: Calculate number of Promoters per region.
  • Step 2: Identify the top region.
  • Step 3: Suggest an action for business growth.
promoter_counts = nps_df[nps_df['Segment'] == 'Promoter'].groupby('Region').size()
top_region = promoter_counts.idxmax()
print(f'Region with most Promoters: {top_region}')
print('Recommendation: Focus new referral campaigns in this region to build on current loyalty.')
Region with most Promoters: North
Recommendation: Focus new referral campaigns in this region to build on current loyalty.

YouTube: Learn More Market Research Analytics#

  • For more hands-on tips, check out YouTubesearch for 'Python market research analytics' and follow along with examples.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.