Mathew K Analytics

Lesson 3 · Market Research Analytics in Python

Quantitative vs Qualitative Research Methods: Key Differences for Market Analysts

In this lesson, we will discover how to apply both quantitative and qualitative approaches to market research and customer analytics. These methods help…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Quantitative vs Qualitative Research Methods in Market Research#

  • In this lesson, we will discover how to apply both quantitative and qualitative approaches to market research and customer analytics.
  • These methods help companies understand what customers do, how they feel, and why they act in certain ways.
  • You will use real customer survey and feedback data to extract actionable business insights.
  • By the end, you will know how to select and apply the right type of analysis for various business questions.
import openml
import pandas as pd
import numpy as np
import warnings
warnings.filterwarnings('ignore')

Key Concepts: Quantitative and Qualitative Research in Market Research#

  • Quantitative research uses structured data like survey ratings, sales numbers, or demographic variables.
  • Qualitative research explores open-ended responses, customer comments, and subjective impressions.
  • Both are crucial: quantitative data shows 'what' is happening; qualitative data helps explain 'why.'
  • Customer analytics relies on collecting the right data and analyzing it using suitable methods.
  • Common mistakes include confusing scales (e.g., 1=bad or 1=good?) and ignoring missing or ambiguous feedback.
# Beginner Example 1: Quantitative survey data (customer satisfaction)
dataset = openml.datasets.get_dataset(42178)
df_quant = dataset.get_data(dataset_format='dataframe')[0]
print(df_quant.shape)
print(df_quant.head(3))
(7043, 20)
   gender  SeniorCitizen Partner Dependents  tenure PhoneService  \
0  Female              0     Yes         No       1           No   
1    Male              0      No         No      34          Yes   
2    Male              0      No         No       2          Yes   

      MultipleLines InternetService OnlineSecurity OnlineBackup  \
0  No phone service             DSL             No          Yes   
1                No             DSL            Yes           No   
2                No             DSL            Yes          Yes   

  DeviceProtection TechSupport StreamingTV StreamingMovies        Contract  \
0               No          No          No              No  Month-to-month   
1              Yes          No          No              No        One year   
2               No          No          No              No  Month-to-month   

  PaperlessBilling     PaymentMethod  MonthlyCharges TotalCharges Churn  
0              Yes  Electronic check           29.85        29.85    No  
1               No      Mailed check           56.95       1889.5    No  
2              Yes      Mailed check           53.85       108.15   Yes  
# Beginner Example 2: Qualitative open-ended feedback
df_qual = pd.DataFrame({'CustomerID':[1,2,3,4,5], 'Feedback':['Great service and friendly staff','Delivery was slow and packaging was poor','Excellent quality, will buy again','Customer support needs improvement','Good value for money']})
print(df_qual.shape)
print(df_qual.head(3))
(5, 2)
   CustomerID                                  Feedback
0           1          Great service and friendly staff
1           2  Delivery was slow and packaging was poor
2           3         Excellent quality, will buy again
# Beginner Example 3: Quantitative NPS survey (synthetic dataset)
np.random.seed(42)
nps_df = pd.DataFrame({
    'CustomerID': range(1, 21),
    'Age': np.random.randint(18, 70, 20),
    'Region': np.random.choice(['North', 'South', 'East', 'West'], 20),
    'NPS_Score': np.random.randint(0,11,20)
})
print(nps_df.head())
   CustomerID  Age Region  NPS_Score
0           1   56   West          2
1           2   69  South          6
2           3   46  South          3
3           4   32  South          8
4           5   60   West          2
# Intermediate Example 1: Calculate mean NPS by region (quantitative segmentation)
nps_region = nps_df.groupby('Region')['NPS_Score'].mean().reset_index()
print(nps_region)
  Region  NPS_Score
0   East   8.666667
1  North   3.600000
2  South   5.833333
3   West   2.666667
# Intermediate Example 2: Simple sentiment keyword count (qualitative analysis)
keywords = ['great','poor','excellent','improvement','good']
for kw in keywords:
    count = df_qual['Feedback'].str.lower().str.count(kw).sum()
    print(f"Keyword '{kw}' appears {count} times.")
Keyword 'great' appears 1 times.
Keyword 'poor' appears 1 times.
Keyword 'excellent' appears 1 times.
Keyword 'improvement' appears 1 times.
Keyword 'good' appears 1 times.
# Intermediate Example 3: Cross-tabulationCustomer satisfaction by gender
df_quant_crosstab = pd.crosstab(df_quant['gender'], df_quant['Churn'])
print(df_quant_crosstab)
Churn     No  Yes
gender           
Female  2549  939
Male    2625  930
# Advanced Example 1: Quantitative and qualitative combinedsentiment by NPS segment
def nps_label(score):
    if score >= 9: return 'Promoter'
    elif score >= 7: return 'Passive'
    else: return 'Detractor'
nps_df['Segment'] = nps_df['NPS_Score'].apply(nps_label)
example_feedback = {
    'Detractor': ['Too slow','Bad packaging','Expensive','Unhelpful support','Rude staff'],
    'Passive': ['Okay service','Fine','Could be better','Average','Nothing special'],
    'Promoter': ['Excellent support','Love the product','Fast delivery','Great quality','Will recommend']
}
nps_df['Feedback'] = nps_df['Segment'].apply(lambda seg: np.random.choice(example_feedback[seg]))
sentiments = nps_df.groupby('Segment')['Feedback'].apply(lambda x: ', '.join(x.head(2))).reset_index()
print(sentiments)
     Segment                              Feedback
0  Detractor  Unhelpful support, Unhelpful support
1    Passive                      Average, Average
2   Promoter    Love the product, Love the product
# Advanced Example 2: Quantitative analysisfrequency of missing survey responses
missing_counts = df_quant.isnull().sum().sort_values(ascending=False)
missing_counts = missing_counts[missing_counts > 0]
print('Variables with missing values:')
print(missing_counts)
Variables with missing values:
Series([], dtype: int64)
# Advanced Example 3: Market response to campaign (quantitative analysis over time)
dataset = openml.datasets.get_dataset(1461)
df_campaign = dataset.get_data(dataset_format='dataframe')[0]
df_campaign.columns = ['age','job','marital','education','default','balance','housing','loan','contact','day','month','duration','campaign','pdays','previous','poutcome','response']
monthly_counts = df_campaign.groupby('month')['response'].value_counts().unstack().fillna(0)
print(monthly_counts)
response      1    2
month               
apr        2355  577
aug        5559  688
dec         114  100
feb        2208  441
jan        1261  142
jul        6268  627
jun        4795  546
mar         229  248
may       12841  925
nov        3567  403
oct         415  323
sep         310  269
# Error Handling 1: Detect missing or ambiguous survey responses
sample_missing = df_quant[df_quant.isnull().any(axis=1)]
print('Rows with missing responses:')
print(sample_missing.head())
Rows with missing responses:
Empty DataFrame
Columns: [gender, SeniorCitizen, Partner, Dependents, tenure, PhoneService, MultipleLines, InternetService, OnlineSecurity, OnlineBackup, DeviceProtection, TechSupport, StreamingTV, StreamingMovies, Contract, PaperlessBilling, PaymentMethod, MonthlyCharges, TotalCharges, Churn]
Index: []
# Error Handling 2: Catch grouping errors (wrong columns in groupby)
try:
    df_quant.groupby('UnknownColumn').size()
except Exception as e:
    print('Grouping failed:', e)
Grouping failed: 'UnknownColumn'
# Error Handling 3: Interpreting NPS scores (Likert confusion)
print('NPS example scores:')
print(nps_df[['CustomerID','NPS_Score','Segment']].head())
print('Remember: Promoters = 9-10, Passives = 7-8, Detractors = 0-6.')
NPS example scores:
   CustomerID  NPS_Score    Segment
0           1          2  Detractor
1           2          6  Detractor
2           3          3  Detractor
3           4          8    Passive
4           5          2  Detractor
Remember: Promoters = 9-10, Passives = 7-8, Detractors = 0-6.

Best Practices and Patterns in Market Research Analytics#

  • Use segmentation to break down metrics by customer type, channel, or geography.
  • Cross-tabulation reveals relationships between demographics and outcomes.
  • Index or score construction turns raw data into easy-to-use business metrics.
  • Trend analysis helps spot seasonality and track changes after business actions.
# Segmentation Example: Average monthly charges by contract type
avg_monthly = df_quant.groupby('Contract')['MonthlyCharges'].mean().reset_index()
print(avg_monthly)
         Contract  MonthlyCharges
0  Month-to-month       66.398490
1        One year       65.048608
2        Two year       60.770413
# Cross-tabulation Pattern: Payment method vs. churn status
pay_churn_ct = pd.crosstab(df_quant['PaymentMethod'], df_quant['Churn'])
print(pay_churn_ct)
Churn                        No   Yes
PaymentMethod                        
Bank transfer (automatic)  1286   258
Credit card (automatic)    1290   232
Electronic check           1294  1071
Mailed check               1304   308
# Index Construction: Create a simple satisfaction index from multiple variables
satisfaction_vars = ['OnlineSecurity','TechSupport','StreamingTV','StreamingMovies']
index_map = { 'Yes': 2, 'No': 1, 'No internet service': 0 }
df_quant['SatisfactionIndex'] = df_quant[satisfaction_vars].replace(index_map).sum(axis=1)
print(df_quant[['SatisfactionIndex']].head())
   SatisfactionIndex
0                  4
1                  5
2                  5
3                  6
4                  4
# Trend Analysis: Monthly customer activity (cohort retention example)
np.random.seed(0)
dates = pd.date_range('2021-01-01', periods=24, freq='ME')
df_cohort = pd.DataFrame({
    'CustomerID': np.random.randint(1000,2000,len(dates)),
    'Signup_Month': dates,
    'Active_Users': np.random.randint(50,300,len(dates))
})
print(df_cohort.head())
print('Average Active Users per Month:', round(df_cohort['Active_Users'].mean(),2))
   CustomerID Signup_Month  Active_Users
0        1684   2021-01-31           138
1        1559   2021-02-28           131
2        1629   2021-03-31           215
3        1192   2021-04-30            75
4        1835   2021-05-31           127
Average Active Users per Month: 181.5
# End-to-End Example: What drives high NPS scores?
grouped = nps_df.groupby('Segment')['Age'].mean().reset_index()
print('Average age by NPS segment:')
print(grouped)
feedback_counts = nps_df.groupby('Segment')['Feedback'].value_counts().groupby(level=0).head(2)
print('Top feedback keywords by segment:')
print(feedback_counts)
Average age by NPS segment:
     Segment        Age
0  Detractor  42.857143
1    Passive  36.250000
2   Promoter  40.000000
Top feedback keywords by segment:
Segment    Feedback         
Detractor  Unhelpful support    6
           Bad packaging        5
Passive    Average              3
           Nothing special      1
Promoter   Love the product     2
Name: count, dtype: int64

Next Steps#

  • You have learned to combine quantitative and qualitative methods for better customer decisions.
  • Keep practicing by applying these techniques to your company's own survey and feedback data.
  • Want more real-world market research walkthroughs? Subscribe to our YouTube channel!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.