Lesson 37 · Market Research Analytics in Python
Measuring Customer Satisfaction with Market Research Data
In this lesson, we will learn how to analyze customer satisfaction survey data for actionable business insights. Customer satisfaction is crucial for…
- CourseMarket Research Analytics in Python
- Lesson37 of 56
- Video25 min
- FormatJupyter notebook · 23 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbMeasuring Customer Satisfaction with Market Research Data#
- In this lesson, we will learn how to analyze customer satisfaction survey data for actionable business insights.
- Customer satisfaction is crucial for building loyalty, improving retention, and informing marketing strategies.
- By the end, you will know how to measure, visualize, and segment satisfaction scores using real datasets.
- We will solve problems relevant to customer analytics and market research teams.
import openml
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns
import warnings
warnings.filterwarnings('ignore')
Understanding Customer Satisfaction Data#
- Most customer satisfaction datasets use structured responses from surveys.
- Columns often include demographics (like age, gender), service indicators, and satisfaction scores (like NPS or Likert ratings).
- Be careful: missing data, incorrectly coded scales, or reversed scales are common mistakes for beginners.
- Always verify what each score means before analyzing, and watch out for leading or biased survey questions.
# BEGINNER EXAMPLE 1: Load a real customer satisfaction survey dataset
dataset = openml.datasets.get_dataset(42178)
df, _, _, _ = dataset.get_data(dataset_format='dataframe')
print(df.shape)
print(df.head(3))
# BEGINNER EXAMPLE 2: Check for missing values in survey responses
missing = df.isnull().sum()
print(missing[missing > 0])
# BEGINNER EXAMPLE 3: Summarize basic demographic information
print(df['gender'].value_counts())
print(df['SeniorCitizen'].value_counts())
# BEGINNER EXAMPLE 4: Visualize satisfaction scores (churn as proxy)
sns.countplot(data=df, x='Churn')
plt.title('Customer Churn as Satisfaction Indicator')
plt.xlabel('Customer Churned?')
plt.ylabel('Count')
plt.show()
# BEGINNER EXAMPLE 5: Calculate and print churn rate
churn_rate = df['Churn'].value_counts(normalize=True)['Yes'] * 100
print(f'Overall churn rate: {churn_rate:.2f}%')
# BEGINNER EXAMPLE 6: Load a synthetic Net Promoter Score (NPS) survey dataset
np.random.seed(42)
nps_df = pd.DataFrame({'CustomerID': range(1,501),
'Age': np.random.randint(18,70,500),
'Region': np.random.choice(['North','South','East','West'],500),
'NPS_Score': np.random.randint(0,11,500)})
print(nps_df.head(3))
# INTERMEDIATE EXAMPLE 1: Visualize the NPS score distribution
sns.histplot(nps_df['NPS_Score'], bins=11, kde=False)
plt.title('NPS Score Distribution')
plt.xlabel('NPS Score (0-10)')
plt.ylabel('Number of Responses')
plt.show()
# INTERMEDIATE EXAMPLE 2: Calculate NPS according to standard formula
total = len(nps_df)
promoters = len(nps_df[nps_df['NPS_Score'] >= 9])
detractors = len(nps_df[nps_df['NPS_Score'] <= 6])
nps = ((promoters - detractors) / total) * 100
print(f'Net Promoter Score (NPS): {nps:.1f}')
# INTERMEDIATE EXAMPLE 3: Segment NPS by region
region_nps = nps_df.groupby('Region').apply(
lambda d: ((len(d[d['NPS_Score'] >= 9]) - len(d[d['NPS_Score'] <= 6])) / len(d)) * 100
)
print(region_nps)
# INTERMEDIATE EXAMPLE 4: Visualize regional NPS differences
region_nps.plot(kind='bar', color='skyblue')
plt.title('Regional Net Promoter Score (NPS)')
plt.xlabel('Region')
plt.ylabel('NPS')
plt.ylim(-100, 100)
plt.show()
# INTERMEDIATE EXAMPLE 5: Load and preview open-ended customer feedback data
feedback_df = pd.DataFrame({
'CustomerID':[1,2,3,4,5],
'Feedback':['Great service and friendly staff',
'Delivery was slow and packaging was poor',
'Excellent quality, will buy again',
'Customer support needs improvement',
'Good value for money']
})
print(feedback_df.head())
# INTERMEDIATE EXAMPLE 6: Simple sentiment identification
feedback_df['Sentiment'] = feedback_df['Feedback'].str.contains('great|excellent|good|friendly|buy again', case=False).map({True: 'Positive', False: 'Negative'})
print(feedback_df[['Feedback', 'Sentiment']])
# ADVANCED EXAMPLE 1: Cross-tabulation between Churn and Internet Service
ct = pd.crosstab(df['InternetService'], df['Churn'], normalize='index') * 100
print(ct)
# ADVANCED EXAMPLE 2: Monthly churn trend analysis
df['SignupMonth'] = (df['tenure'] // 1).astype(int)
monthly_churn = df.groupby('SignupMonth')['Churn'].apply(lambda x: (x=='Yes').mean() * 100)
monthly_churn.plot(kind='line', marker='o')
plt.title('Monthly Churn Rate Trend')
plt.xlabel('Month Since Signup')
plt.ylabel('Churn Rate (%)')
plt.show()
# ADVANCED EXAMPLE 3: Create a Customer Satisfaction Index
satisfaction_index = df[['MonthlyCharges', 'tenure']].apply(lambda x: (x['MonthlyCharges']/x['tenure']) if x['tenure'] else x['MonthlyCharges'], axis=1)
df['SatisfactionIndex'] = satisfaction_index
print(df[['MonthlyCharges', 'tenure', 'SatisfactionIndex']].head(3))
# Median Imputation for Numeric Columns
df_filled = df.copy()
num_cols = ['TotalCharges', 'MonthlyCharges']
df_filled[num_cols] = df_filled[num_cols].apply(
pd.to_numeric, errors='coerce'
)
df_filled[num_cols] = df_filled[num_cols].fillna(
df_filled[num_cols].median()
)
print(df_filled[num_cols].isnull().sum())
# ERROR HANDLING 2: Guard against grouping errors
try:
result = df.groupby('Gender').size()
except KeyError:
print('Check column spelling: Should be gender (lowercase), not Gender.')
# ERROR HANDLING 3: Watch out for misinterpreted NPS scale
def recode_nps(score):
try:
s = int(score)
if s > 10 or s < 0:
return np.nan
else:
return s
except:
return np.nan
nps_df['NPS_Score'] = nps_df['NPS_Score'].apply(recode_nps)
print(nps_df['NPS_Score'].isnull().sum())
Best Practices for Customer Satisfaction Analytics#
- Segment and compare satisfaction by key demographics or behaviors.
- Use cross-tabulation for quick two-way comparisons (e.g., product vs. churn).
- Always standardize rating scales so everyone is measured on the same terms.
- Trend analysis over time highlights critical moments in the customer journey.
- Build indices when simple single-question scores are not enough.
# PATTERN EXAMPLE 1: Customer segmentation by contract type
seg_ct = pd.crosstab(df['Contract'], df['Churn'], normalize='index') * 100
print(seg_ct)
# PATTERN EXAMPLE 2: Index construction for advanced benchmarking
df['ChargeIndex'] = df['MonthlyCharges']/df['MonthlyCharges'].max()
df['TenureIndex'] = df['tenure']/df['tenure'].max()
print(df[['ChargeIndex', 'TenureIndex']].head(3))
# FULL MARKET RESEARCH PROBLEM: From survey data to a clear recommendation
dataset = openml.datasets.get_dataset(42178)
df, _, _, _ = dataset.get_data(dataset_format='dataframe')
df['MonthlyCharges'] = df['MonthlyCharges'].fillna(df['MonthlyCharges'].median())
nps_proxy = (df['Churn'].map({'No': 10, 'Yes': 0}))
high_value = df['MonthlyCharges'] > df['MonthlyCharges'].median()
promoters_pct = (nps_proxy[high_value] >= 9).sum() / high_value.sum() * 100
print(f'Percentage of high-value customers who are promoters: {promoters_pct:.2f}%')
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



