Lesson 38 · Market Research Analytics in Python
Net Promoter Score Analysis in Python for Market Research
In this lesson, we solve a real-world market research problem using customer survey data. Net Promoter Score (NPS) helps businesses understand customer…
- CourseMarket Research Analytics in Python
- Lesson38 of 56
- Video20 min
- FormatJupyter notebook · 20 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbNet Promoter Score (NPS) Analysis#
- In this lesson, we solve a real-world market research problem using customer survey data.
- Net Promoter Score (NPS) helps businesses understand customer loyalty and word-of-mouth potential.
- You will learn to calculate, segment, and analyze NPS results for actionable business insights.
- We start from raw survey data and will produce clear, easy-to-use NPS insights for decision-making.
- Let us explore why NPS matters and how to avoid common survey analysis mistakes.
import pandas as pd
import numpy as np
import warnings
warnings.filterwarnings('ignore')
Understanding NPS survey data#
- NPS surveys ask customers how likely they are to recommend a company on a 0-10 scale.
- Responses group into: 0-6 (Detractors), 7-8 (Passives), and 9-10 (Promoters).
- NPS = % Promoters minus % Detractors.
- Surveys also often include basic customer info such as age, region, or product used.
- Beginners often misinterpret scales, forget to segment NPS by customer groups, or skip checking for missing data.
# Beginner Example 1: Load a synthetic NPS survey dataset
np.random.seed(42)
df = pd.DataFrame({
'CustomerID': range(1, 501),
'Age': np.random.randint(18, 70, 500),
'Region': np.random.choice(['North', 'South', 'East', 'West'], 500),
'NPS_Score': np.random.randint(0, 11, 500)
})
print('Shape:', df.shape)
print(df.head(3))
# Beginner Example 2: Basic NPS calculation
promoters = (df['NPS_Score'] >= 9).sum()
detractors = (df['NPS_Score'] <= 6).sum()
n_total = len(df)
nps = (promoters / n_total) * 100 - (detractors / n_total) * 100
print(f'NPS Score: {nps:.1f}')
# Beginner Example 3: Visualizing NPS score distribution
import matplotlib.pyplot as plt
fig, ax = plt.subplots(figsize=(8,4))
df['NPS_Score'].hist(bins=11, color='steelblue', edgecolor='black', ax=ax)
ax.set_title('Distribution of NPS Scores')
ax.set_xlabel('NPS Score (0-10)')
ax.set_ylabel('Number of Responses')
plt.tight_layout()
plt.show()
# Beginner Example 4: Count each NPS response group
df['Group'] = pd.cut(df['NPS_Score'], bins=[-1,6,8,10], labels=['Detractor','Passive','Promoter'])
group_counts = df['Group'].value_counts().sort_index()
print(group_counts)
# Beginner Example 5: Segmenting NPS by region
nps_by_region = df.groupby('Region').apply(lambda x: ((x['NPS_Score'] >= 9).mean() - (x['NPS_Score'] <= 6).mean()) * 100)
print(nps_by_region)
# Intermediate Example 1: Explore NPS by age group
df['AgeGroup'] = pd.cut(df['Age'], bins=[17,29,39,49,59,100], labels=['18-29','30-39','40-49','50-59','60+'])
nps_age = df.groupby('AgeGroup').apply(lambda x: ((x['NPS_Score'] >= 9).mean() - (x['NPS_Score'] <= 6).mean()) * 100)
print(nps_age)
# Intermediate Example 2: Cross-tabulation of NPS group by region
crosstab = pd.crosstab(df['Region'], df['Group'], normalize='index').round(2)
print(crosstab)
# Intermediate Example 3: Plot NPS by region
nps_by_region.plot(kind='bar', color='seagreen', title='NPS by Region')
plt.ylabel('NPS Score')
plt.xlabel('Region')
plt.tight_layout()
plt.show()
# Intermediate Example 4: Calculate NPS and response counts with aggregation
agg = df.groupby('Region').agg(
n_promoters = ('NPS_Score', lambda x: (x >= 9).sum()),
n_detractors = ('NPS_Score', lambda x: (x <= 6).sum()),
total = ('NPS_Score', 'count')
)
agg['NPS'] = (agg['n_promoters'] / agg['total'] - agg['n_detractors'] / agg['total']) * 100
print(agg[['NPS']])
# Intermediate Example 5: Checking for missing responses
missing_nps = df['NPS_Score'].isnull().sum()
print(f'Missing NPS responses: {missing_nps}')
# Intermediate Example 6: Remove missing NPS scores from analysis
filtered_df = df.dropna(subset=['NPS_Score'])
print(f'Filtered dataset: {filtered_df.shape[0]} responses (from {df.shape[0]})')
# Advanced Example 1: Calculate NPS confidence interval
from statsmodels.stats.proportion import proportion_confint
n_prom = (df['NPS_Score'] >= 9).sum()
n_det = (df['NPS_Score'] <= 6).sum()
n = len(df)
prop_prom = n_prom / n
prop_det = n_det / n
ci_prom = proportion_confint(n_prom, n, method='wilson')
ci_det = proportion_confint(n_det, n, method='wilson')
nps_ci = [100 * (ci_prom[0] - ci_det[1]), 100 * (ci_prom[1] - ci_det[0])]
print(f'NPS Confidence Interval: [{nps_ci[0]:.1f}, {nps_ci[1]:.1f}]')
# Advanced Example 2: Trend analysis using a time-based NPS field
df['SurveyMonth'] = np.random.choice(pd.date_range('2021-01-01', periods=6, freq='M'), size=len(df))
trend = df.groupby(df['SurveyMonth'].dt.to_period('M')).apply(lambda x: ((x['NPS_Score'] >= 9).mean() - (x['NPS_Score'] <= 6).mean()) * 100)
trend.plot(marker='o', title='NPS Trend Over Time')
plt.ylabel('NPS Score')
plt.xlabel('Month')
plt.show()
# Advanced Example 3: Merge NPS data with open-ended feedback
feedback_df = pd.DataFrame({'CustomerID':[1,2,3,4,5], 'Feedback':[
'Great service and friendly staff',
'Delivery was slow and packaging was poor',
'Excellent quality, will buy again',
'Customer support needs improvement',
'Good value for money'
]})
merged = pd.merge(df, feedback_df, on='CustomerID', how='left')
print(merged[['CustomerID','NPS_Score','Group','Feedback']].head(7))
# Error Handling 1: What if a user enters a score outside 0-10?
invalid = (df['NPS_Score'] < 0) | (df['NPS_Score'] > 10)
print(f'Number of invalid NPS scores: {invalid.sum()}')
# Error Handling 2: Handle missing or bad region values during group calculations
if df['Region'].isnull().any():
print('Missing regions found. Filling with Unknown.')
df['Region'] = df['Region'].fillna('Unknown')
# Error Handling 3: Avoid integer division mistakes in NPS calculation
# (Python 2 would floor divide; Python 3 behavior is safer but check anyway)
sample_prom = 5
sample_det = 3
sample_total = 10
nps_int = (sample_prom // sample_total - sample_det // sample_total) * 100
nps_float = (sample_prom / sample_total - sample_det / sample_total) * 100
print('NPS using integer division:', nps_int)
print('NPS using float division:', nps_float)
Best Practices: NPS segmentation, cross-tabulation, and scoring patterns#
- Always segment NPS by customer demographics or business area for actionable detail.
- Use cross-tabulation to spot patterns, such as which product or region drives high NPS.
- Index construction (such as NPS) turns survey ratings into a business metric.
- Analyze trends over time to validate business improvements or flag emerging risks.
- Document every calculation for transparency in market research reporting.
# End-to-End Example: From raw NPS data to clear business recommendation
# 1. Clean the data (remove invalid NPS scores)
cleaned = df[(df['NPS_Score'] >= 0) & (df['NPS_Score'] <= 10)]
n_total = len(cleaned)
# 2. Calculate overall company NPS
nps_score = ((cleaned['NPS_Score'] >= 9).mean() - (cleaned['NPS_Score'] <= 6).mean()) * 100
# 3. Segment by region to identify the lowest and highest NPS
nps_by_region = cleaned.groupby('Region').apply(lambda x: ((x['NPS_Score'] >= 9).mean() - (x['NPS_Score'] <= 6).mean()) * 100)
best_region = nps_by_region.idxmax()
worst_region = nps_by_region.idxmin()
# 4. Output insights and recommendation
print(f'Company overall NPS: {nps_score:.1f}')
print(nps_by_region)
print(f'Highest NPS region: {best_region}')
print(f'Lowest NPS region: {worst_region}')
if nps_by_region[worst_region] < 0:
print(f'Recommendation: Focus retention programs in {worst_region}.')
else:
print(f'All regions positive. Continue to invest in {best_region} for best practices.')
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



