Lesson 13 · Market Research Analytics in Python
Likert Scale and Rating Question Analysis in Python for Market Research
Are customers satisfied with our services? How do we measure customer sentiment using survey ratings? What business insights can we gain from consistent…
- CourseMarket Research Analytics in Python
- Lesson13 of 56
- Video21 min
- FormatJupyter notebook · 20 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbLikert Scale and Rating Question Analysis#
Are customers satisfied with our services?
How do we measure customer sentiment using survey ratings?
What business insights can we gain from consistent survey analysis?
In this lesson, we will learn how to interpret and analyze Likert scale data and rating questions, turning raw survey responses into actionable insights.
These skills are vital for businesses to improve products, services, and customer experience.
import pandas as pd
import numpy as np
import openml
import warnings
warnings.filterwarnings('ignore')
What are Likert scales and rating questions?#
Surveys often include rating questions (1-5, Agree-Disagree) called Likert scales.
Each response reflects the attitude or satisfaction of the customer.
Demographics such as age or gender may also be included for segmentation.
Beginners often:
- Treat Likert scales as purely numeric scores without checking scale direction.
- Ignore missing or non-response values.
- Fail to distinguish between ordinal and nominal data, leading to incorrect analyses.
- Ignore missing or non-response values.
- Treat Likert scales as purely numeric scores without checking scale direction.
# Example 1: Load a customer satisfaction dataset with Likert scale ratings
dataset = openml.datasets.get_dataset(42178)
df, _, _, _ = dataset.get_data(dataset_format='dataframe')
print(df.shape)
print(df.head(3))
Example 2: What do survey columns mean?#
- Common columns in customer satisfaction datasets:
- Demographics: gender, age, SeniorCitizen, region
- Service usage: PhoneService, InternetService, OnlineSecurity, etc.
- Survey responses: Often encoded as 0 (No), 1 (Yes), or 1-5 for ratings
- Target variable: Churn (has the customer left?)
- Survey responses: Often encoded as 0 (No), 1 (Yes), or 1-5 for ratings
- Service usage: PhoneService, InternetService, OnlineSecurity, etc.
- Demographics: gender, age, SeniorCitizen, region
# Example 3: Count responses for a Likert-scale question (e.g., OnlineSecurity)
likert_counts = df['OnlineSecurity'].value_counts(dropna=False)
print(likert_counts)
# Example 4: Calculate the percentage distribution for a Likert-scale question
likert_percent = df['OnlineSecurity'].value_counts(normalize=True, dropna=False) * 100
print(likert_percent.round(2))
# Example 5: Visualize Likert-scale ratings (bar plot)
import matplotlib.pyplot as plt
likert_percent.plot(kind='bar', color='skyblue')
plt.title('Online Security - Survey Responses (%)')
plt.ylabel('Percentage')
plt.xlabel('Response')
plt.show()
# Example 6: Calculate mean satisfaction rating for a numeric scale (e.g., MonthlyCharges)
mean_charge = df['MonthlyCharges'].mean()
print(f'Average Monthly Charges: ${mean_charge:.2f}')
# Example 7: Segment Likert-scale ratings by customer demographics (e.g., gender)
grouped = df.groupby('gender')['OnlineSecurity'].value_counts(normalize=True).unstack().fillna(0) * 100
print(grouped.round(1))
# Example 8: Cross-tabulate two categorical variables (OnlineSecurity and Churn)
ct = pd.crosstab(df['OnlineSecurity'], df['Churn'], normalize='columns') * 100
print(ct.round(1))
# Example 9: Create a satisfaction index from several Likert-scale columns
cols = ['OnlineSecurity', 'TechSupport', 'DeviceProtection']
df['satisfaction_index'] = df[cols].applymap(lambda x: 1 if x == 'Yes' else 0).mean(axis=1)
print(df[['satisfaction_index']].head())
# Example 10: Trend analysis for satisfaction index (by tenure in years)
df['tenure_years'] = (df['tenure'] // 12).astype(int)
trend = df.groupby('tenure_years')['satisfaction_index'].mean()
print(trend)
# Example 11: Advanced - Filter for survey non-response / missing data
missing = df[cols].isnull().sum()
print('Missing responses in key satisfaction columns:')
print(missing)
# Example 12: Advanced - Handle missing values by filling with 'No Response'
df_filled = df.copy()
df_filled[cols] = df_filled[cols].fillna('No Response')
print(df_filled[cols].head())
# Example 13: Advanced - Misinterpreted Likert scales: reverse coding
df['Feedback_Quality'] = np.random.choice([1,2,3,4,5], size=len(df))
df['Quality_reversed'] = 6 - df['Feedback_Quality']
print(df[['Feedback_Quality','Quality_reversed']].head())
# Example 14: Error Handling - Wrong groupby column name
try:
df.groupby('gendre')['OnlineSecurity'].mean()
except Exception as e:
print('Error:', e)
# Example 15: Error Handling - Interpreting NPS scores as Likert scale
nps_df = pd.DataFrame({'NPS_Score': np.random.randint(0, 11, 100)})
print(nps_df['NPS_Score'].describe())
print('Mean NPS: Not directly interpretable as Net Promoter Score, use the correct classification!')
# Example 16: Best Practice - Calculate NPS breakdown
nps_df['Type'] = np.where(nps_df['NPS_Score'] >= 9, 'Promoter',
np.where(nps_df['NPS_Score'] <= 6, 'Detractor', 'Passive'))
nps_counts = nps_df['Type'].value_counts(normalize=True) * 100
print('NPS Classification (%):')
print(nps_counts.round(1))
Segmentation, cross-tabs, and trend analysis#
- Key analytics patterns for market research:
- Segmentation: split responses by demographic or behavioral groups
- Cross-tabulation: relate two categorical variables to show associations
- Index or score construction: combine multiple responses into one strong metric
Trend analysis: track changes across time or tenure
Applying these patterns helps surface business-critical insights.
- Index or score construction: combine multiple responses into one strong metric
- Cross-tabulation: relate two categorical variables to show associations
- Segmentation: split responses by demographic or behavioral groups
# Example 17: Best Practice - Segmentation by SeniorCitizen and satisfaction
seg = df.groupby('SeniorCitizen')['satisfaction_index'].mean()
print('Average satisfaction index by senior citizen status:')
print(seg.round(3))
# Example 18: Best Practice - Multi-way cross-tab (Churn by InternetService and satisfaction)
df['satisfaction_bin'] = pd.cut(df['satisfaction_index'], bins=[-0.1,0.33,0.66,1], labels=['Low','Medium','High'])
crosstab = pd.crosstab([df['InternetService'], df['satisfaction_bin']], df['Churn'], normalize='columns') * 100
print(crosstab.round(1))
# Example 19: Trend chart - satisfaction index by tenure
trend.plot(marker='o')
plt.xlabel('Tenure (years)')
plt.ylabel('Avg Satisfaction Index')
plt.title('Satisfaction Trend by Customer Tenure')
plt.show()
# Example 20: Tiny end-to-end market research workflow summary
def survey_analysis_workflow(df):
# Step 1: Check missing values
missing = df[cols].isnull().sum().sum()
# Step 2: Compute satisfaction index
satisfaction = df[cols].applymap(lambda x: 1 if x=='Yes' else 0).mean(axis=1)
# Step 3: Split by churn
churned = satisfaction[df['Churn']=='Yes'].mean()
not_churned = satisfaction[df['Churn']=='No'].mean()
print(f'Total missing responses: {missing}')
print(f'Churned Customers Avg Satisfaction: {churned:.3f}')
print(f'Non-Churned Customers Avg Satisfaction: {not_churned:.3f}')
if churned < not_churned:
print('Insight: Lower satisfaction is linked to higher churn. Prioritize improvements!')
else:
print('No strong satisfaction-churn link found.')
survey_analysis_workflow(df)
Summary and what to try next#
- You have learned to:
- Analyze and visualize Likert scale and rating survey data
- Handle missing values and avoid common pitfalls
- Segment, cross-tabulate, and construct useful analytics indices
Connect survey insights to business recommendations
For more hands-on survey analytics, review video case studies on YouTube!
- Segment, cross-tabulate, and construct useful analytics indices
- Handle missing values and avoid common pitfalls
- Analyze and visualize Likert scale and rating survey data
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



