Lesson 8 · Market Research Analytics in Python
Numerical Analysis with NumPy for Market Data Training
Learn how to use NumPy to analyze real-world market research and customer analytics datasets. We focus on basic and advanced numerical techniques to uncover…
- CourseMarket Research Analytics in Python
- Lesson8 of 56
- Video23 min
- FormatJupyter notebook · 25 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbNumerical Analysis with NumPy for Market Data#
- Learn how to use NumPy to analyze real-world market research and customer analytics datasets.
- We focus on basic and advanced numerical techniques to uncover business insights.
- You will analyze customer satisfaction, survey scores, and campaign performance using Python.
- These skills help companies make data-driven decisions, improve products, and understand market trends.
- By the end, you will summarize, aggregate, and visualize customer data to answer real business questions.
import pandas as pd
import numpy as np
import openml
import warnings
warnings.filterwarnings('ignore')
Understanding Market Data and Surveys#
- Market research data can include customer responses, satisfaction scores, or sales information.
- Datasets often have columns for demographic data, survey scores, and purchase behavior.
- Poor data handling may include misclassifying survey scales, missing values, or incorrect aggregations.
- Beginners often forget to validate data types or miss handling negative responses correctly.
# Load the Customer Satisfaction dataset from OpenML
dataset = openml.datasets.get_dataset(42178)
df = dataset.get_data(dataset_format='dataframe')[0]
print(df.shape)
print(df.head(3))
# Beginner: Calculate the mean tenure of customers
mean_tenure = df['tenure'].mean()
print(f'Average customer tenure: {mean_tenure:.2f} months')
# Beginner: Find the proportion of customers who churned
churn_rate = (df['Churn'] == 'Yes').mean()
print(f'Churn rate: {churn_rate:.2%}')
# Beginner: Count male and female customers using NumPy
vals, counts = np.unique(df['gender'], return_counts=True)
for v, c in zip(vals, counts):
print(f'{v}: {c}')
# Intermediate: Create a NumPy array of MonthlyCharges and calculate statistics
charges = df['MonthlyCharges'].values
print('Mean:', np.mean(charges))
print('Std Dev:', np.std(charges))
print('Min:', np.min(charges), 'Max:', np.max(charges))
# Intermediate: Find correlation between tenure and monthly charges
corr = np.corrcoef(df['tenure'], df['MonthlyCharges'])[0,1]
print(f'Correlation between tenure and monthly charges: {corr:.2f}')
# Intermediate: Calculate average TotalCharges for churned vs. non-churned customers
df['TotalCharges'] = pd.to_numeric(df['TotalCharges'], errors='coerce')
churned = df[df['Churn'] == 'Yes']['TotalCharges']
not_churned = df[df['Churn'] == 'No']['TotalCharges']
print('Avg TotalCharges (Churned):', np.nanmean(churned))
print('Avg TotalCharges (Not Churned):', np.nanmean(not_churned))
# Advanced: Use NumPy to create bins of MonthlyCharges
bins = np.arange(0, df['MonthlyCharges'].max()+10, 10)
labels = [f'{int(b)}-{int(b+10)}' for b in bins[:-1]]
df['ChargeBin'] = pd.cut(df['MonthlyCharges'], bins=bins, labels=labels)
charge_counts = df['ChargeBin'].value_counts().sort_index()
print(charge_counts)
# Advanced: Calculate the churn rate for each charge bin
churn_by_bin = df.groupby('ChargeBin')['Churn'].apply(lambda x: (x == 'Yes').mean())
print(churn_by_bin)
# Advanced: Calculate average tenure for each contract type using NumPy
contract_types = df['Contract'].unique()
for ctype in contract_types:
avg_tenure = np.mean(df[df['Contract']==ctype]['tenure'])
print(f'{ctype}: {avg_tenure:.1f} months')
# Load a synthetic NPS survey dataset for practice
np.random.seed(42)
df_nps = pd.DataFrame({
'CustomerID': range(1,501),
'Age': np.random.randint(18,70,500),
'Region': np.random.choice(['North','South','East','West'],500),
'NPS_Score': np.random.randint(0,11,500)
})
print(df_nps.head(3))
# Beginner: Calculate mean NPS score overall
print('Mean NPS score:', np.mean(df_nps['NPS_Score']))
# Intermediate: Calculate the percentage of Promoters, Passives, and Detractors
promoters = (df_nps['NPS_Score'] >= 9).mean()
passives = ((df_nps['NPS_Score'] >= 7) & (df_nps['NPS_Score'] <= 8)).mean()
detractors = (df_nps['NPS_Score'] <= 6).mean()
print(f'Promoters: {promoters:.2%}')
print(f'Passives: {passives:.2%}')
print(f'Detractors: {detractors:.2%}')
# Intermediate: Calculate NPS by Region using NumPy
regions = df_nps['Region'].unique()
for region in regions:
scores = df_nps[df_nps['Region']==region]['NPS_Score']
promoters = np.sum(scores >= 9)
detractors = np.sum(scores <= 6)
n = len(scores)
nps = (promoters - detractors) / n * 100
print(f'Region: {region}, NPS: {nps:.1f}')
# Advanced: Detect missing values in the NPS dataset
missing = df_nps.isnull().sum()
print('Missing values per column:')
print(missing)
# Error Handling: What happens on missing survey responses?
df_nps_missing = df_nps.copy()
df_nps_missing.loc[0:4, 'NPS_Score'] = np.nan
try:
mean_nps = np.mean(df_nps_missing['NPS_Score'])
print('Mean NPS with missing:', mean_nps)
except Exception as e:
print('Error:', e)
# Error Handling: Clean and impute (fill) missing values with the mean score
df_nps_filled = df_nps_missing.copy()
mean_score = np.nanmean(df_nps_filled['NPS_Score'])
df_nps_filled['NPS_Score'] = df_nps_filled['NPS_Score'].fillna(mean_score)
print(df_nps_filled['NPS_Score'].head(6))
# Error Handling: Incorrect aggregation (mean instead of sum for sales)
sales = np.array([20, 30, 50, np.nan, 10])
try:
print('Wrong total sales:', np.mean(sales))
print('Correct total sales:', np.nansum(sales))
except Exception as e:
print('Error:', e)
# Best Practice: Segment customers by tenure group
bins = [0, 12, 24, 36, 48, 60, 72]
labels = ['<1y','1-2y','2-3y','3-4y','4-5y','5y+']
df['TenureGroup'] = pd.cut(df['tenure'], bins=bins, labels=labels, right=False)
print(df['TenureGroup'].value_counts().sort_index())
# Best Practice: Cross-tabulate churn by tenure group
ctab = pd.crosstab(df['TenureGroup'], df['Churn'], normalize='index')
print(ctab)
# Best Practice: Customer Retention Index (CRI) calculation
df['Retained'] = (df['Churn'] == 'No').astype(int)
cri = df.groupby('TenureGroup')['Retained'].mean() * 100
cri = cri.round(1)
print('Customer Retention Index by tenure group (%):')
print(cri)
# Best Practice: Monthly trend analysis of NPS score
df_nps['Month'] = np.random.choice(range(1,13), size=len(df_nps))
monthly_nps = df_nps.groupby('Month')['NPS_Score'].mean()
print('Average NPS score per month:')
print(monthly_nps.sort_index())
# End-to-end Example: Recommend action by identifying highest-churn segment
churn_by_segment = df.groupby('TenureGroup')['Churn'].apply(lambda x: (x == 'Yes').mean()).sort_values(ascending=False)
print('Churn rate by tenure group:')
print(churn_by_segment)
worst_segment = churn_by_segment.idxmax()
print(f'Recommend immediate retention offers for: {worst_segment}')
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



