Lesson 21 · Market Research Analytics in Python
Descriptive Statistics for Market Research Using Python | Comprehensive Analysis Guide
In this lesson, we will learn how to use descriptive statistics to understand customer and market research data. Descriptive statistics help businesses…
- CourseMarket Research Analytics in Python
- Lesson21 of 56
- Video26 min
- FormatJupyter notebook · 23 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbDescriptive Statistics for Market Research#
- In this lesson, we will learn how to use descriptive statistics to understand customer and market research data.
- Descriptive statistics help businesses summarize, visualize, and interpret key patterns in survey responses and customer behavior.
- These insights help businesses make informed decisions about products, services, and marketing campaigns.
- By the end, you will be able to describe customer ratings, NPS, and behaviors using Python.
import pandas as pd
import numpy as np
import openml
import warnings
warnings.filterwarnings('ignore')
Understanding Market Research Data#
- Market research data may come from customer satisfaction surveys, NPS surveys, or tracking online retail behavior.
- Data is often structured as rows of responses, with columns for demographics, ratings, and text feedback.
- Each row usually represents one customer or one transaction.
- Beginners often forget to check for missing values, misunderstand categorical variables, or incorrectly interpret survey scales like NPS.
- Always review your data types and check for any unusual or unexpected values before analysis.
# Beginner Example 1: Load an NPS survey dataset (synthetic, reproducible)
np.random.seed(42)
df_nps = pd.DataFrame({
'CustomerID': range(1,501),
'Age': np.random.randint(18,70,500),
'Region': np.random.choice(['North','South','East','West'],500),
'NPS_Score': np.random.randint(0,11,500)
})
print(df_nps.shape)
print(df_nps.head(3))
# Beginner Example 2: Calculate the mean NPS score
mean_nps = df_nps['NPS_Score'].mean()
print(f'Mean NPS score: {mean_nps:.2f}')
# Beginner Example 3: Find the minimum and maximum NPS scores
min_nps = df_nps['NPS_Score'].min()
max_nps = df_nps['NPS_Score'].max()
print(f'Minimum NPS score: {min_nps}')
print(f'Maximum NPS score: {max_nps}')
# Beginner Example 4: Count NPS promoters, passives, detractors
nps_bins = pd.cut(df_nps['NPS_Score'], bins=[-0.1,6,8,10], labels=['Detractor','Passive','Promoter'])
counts = nps_bins.value_counts()
print(counts)
# Beginner Example 5: Calculate region-wise average NPS
region_avg = df_nps.groupby('Region')['NPS_Score'].mean()
print(region_avg)
# Intermediate Example 1: Load a customer satisfaction dataset from OpenML
dataset = openml.datasets.get_dataset(42178)
df_csat, _, _, _ = dataset.get_data(dataset_format='dataframe')
print(df_csat.shape)
print(df_csat.head(3))
# Intermediate Example 2: Calculate overall monthly charges statistics
monthly_desc = df_csat['MonthlyCharges'].describe()
print(monthly_desc)
# Intermediate Example 3: Gender segmentation of monthly charges
gender_means = df_csat.groupby('gender')['MonthlyCharges'].mean()
print(gender_means)
# Intermediate Example 4: Churn rate calculation
churn_counts = df_csat['Churn'].value_counts(normalize=True)
print(f'Churn rate:\n{churn_counts}')
# Intermediate Example 5: Cross-tabulation - Payment Method vs Churn
payment_churn_crosstab = pd.crosstab(df_csat['PaymentMethod'], df_csat['Churn'])
print(payment_churn_crosstab)
# Advanced Example 1: Load the Online Retail Behavior dataset from UCI
url = 'https://archive.ics.uci.edu/ml/machine-learning-databases/00502/online_retail_II.xlsx'
df_retail = pd.read_excel(url, sheet_name='Year 2010-2011')
df_retail['InvoiceDate'] = pd.to_datetime(df_retail['InvoiceDate'])
print(df_retail.shape)
print(df_retail.head(3))
# Advanced Example 2: Calculate average order value (AOV) by country
df_retail['OrderValue'] = df_retail['Quantity'] * df_retail['Price']
aov_country = df_retail.groupby('Country')['OrderValue'].mean().sort_values(ascending=False)
print(aov_country.head(5))
# Advanced Example 3: Identify top products by revenue
product_revenue = df_retail.groupby('Description')['OrderValue'].sum().sort_values(ascending=False)
print(product_revenue.head(5))
# Advanced Example 4: Trend analysis of monthly sales
df_retail['Month'] = df_retail['InvoiceDate'].dt.to_period('M')
monthly_sales = df_retail.groupby('Month')['OrderValue'].sum()
print(monthly_sales.tail(6))
# Error Handling Example 1: Detect missing survey responses
missing_nps = df_nps.isnull().sum()
print(f'Missing values in NPS dataset:\n{missing_nps}')
# Error Handling Example 2: Guard against incorrect groupings (e.g., misspelled region names)
df_nps['Region'] = df_nps['Region'].str.title()
unique_regions = df_nps['Region'].unique()
print(f'Unique region names: {unique_regions}')
# Error Handling Example 3: Avoid misinterpreting the NPS scale
if df_nps['NPS_Score'].max() > 10 or df_nps['NPS_Score'].min() < 0:
print('Warning: NPS scores should be between 0 and 10! Check your data.')
else:
print('All NPS scores are within the correct range.')
Best Practices and Market Research Patterns#
- Always segment customers by demographics, behavior, or engagement for clearer insights.
- Use cross-tabulation to find relationships and dependencies between survey questions or attributes.
- Build customer indices or scores by combining multiple survey questions, such as satisfaction and loyalty.
- Analyze trends to capture seasonality, campaign effects, or signals of decline and growth.
- Document every transformation or calculation so that insights can be reproduced easily.
# Pattern Example: Create a composite customer satisfaction score
df_csat['SatisfactionIndex'] = (df_csat['MonthlyCharges'].rank(pct=True) +
(1 - df_csat['tenure'].rank(pct=True))) / 2
print(df_csat[['MonthlyCharges','tenure','SatisfactionIndex']].head(3))
# Pattern Example: Cross-tabulate NPS segment by age group
age_bins = pd.cut(df_nps['Age'], bins=[17,25,35,50,70], labels=['18-25','26-35','36-50','51-70'])
nps_age_crosstab = pd.crosstab(age_bins, nps_bins)
print(nps_age_crosstab)
# Pattern Example: Trend in average NPS score by region
df_nps['ResponseMonth'] = np.random.choice(pd.date_range('2023-01-01', periods=12, freq='MS'), df_nps.shape[0])
monthly_nps_region = df_nps.groupby([df_nps['ResponseMonth'].dt.to_period('M'),'Region'])['NPS_Score'].mean().unstack()
print(monthly_nps_region.tail(6))
# Tiny End-to-End Example: From NPS survey to recommendation
overall_nps = (counts['Promoter'] - counts['Detractor']) / len(df_nps) * 100
print(f'Your Net Promoter Score is: {overall_nps:.1f}')
if overall_nps < 0:
print('Your NPS is negative. Business should urgently investigate negative customer experiences.')
elif overall_nps < 30:
print('NPS is low. Consider targeting detractors with win-back offers and collecting further feedback.')
elif overall_nps < 70:
print('NPS is moderate. Maintain current strengths but look for ways to create more promoters.')
else:
print('NPS is excellent. Promote your high Net Promoter Score in marketing!')
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



