Mathew K Analytics

Lesson 13 · Market Research Analytics in Python

Likert Scale and Rating Question Analysis in Python for Market Research

Are customers satisfied with our services? How do we measure customer sentiment using survey ratings? What business insights can we gain from consistent…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Likert Scale and Rating Question Analysis#

  • Are customers satisfied with our services?

  • How do we measure customer sentiment using survey ratings?

  • What business insights can we gain from consistent survey analysis?

  • In this lesson, we will learn how to interpret and analyze Likert scale data and rating questions, turning raw survey responses into actionable insights.

  • These skills are vital for businesses to improve products, services, and customer experience.

import pandas as pd
import numpy as np
import openml
import warnings
warnings.filterwarnings('ignore')

What are Likert scales and rating questions?#

  • Surveys often include rating questions (1-5, Agree-Disagree) called Likert scales.

  • Each response reflects the attitude or satisfaction of the customer.

  • Demographics such as age or gender may also be included for segmentation.

  • Beginners often:

    • Treat Likert scales as purely numeric scores without checking scale direction.
      • Ignore missing or non-response values.
        • Fail to distinguish between ordinal and nominal data, leading to incorrect analyses.
# Example 1: Load a customer satisfaction dataset with Likert scale ratings
dataset = openml.datasets.get_dataset(42178)
df, _, _, _ = dataset.get_data(dataset_format='dataframe')
print(df.shape)
print(df.head(3))
(7043, 20)
   gender  SeniorCitizen Partner Dependents  tenure PhoneService  \
0  Female              0     Yes         No       1           No   
1    Male              0      No         No      34          Yes   
2    Male              0      No         No       2          Yes   

      MultipleLines InternetService OnlineSecurity OnlineBackup  \
0  No phone service             DSL             No          Yes   
1                No             DSL            Yes           No   
2                No             DSL            Yes          Yes   

  DeviceProtection TechSupport StreamingTV StreamingMovies        Contract  \
0               No          No          No              No  Month-to-month   
1              Yes          No          No              No        One year   
2               No          No          No              No  Month-to-month   

  PaperlessBilling     PaymentMethod  MonthlyCharges TotalCharges Churn  
0              Yes  Electronic check           29.85        29.85    No  
1               No      Mailed check           56.95       1889.5    No  
2              Yes      Mailed check           53.85       108.15   Yes  

Example 2: What do survey columns mean?#

  • Common columns in customer satisfaction datasets:
    • Demographics: gender, age, SeniorCitizen, region
      • Service usage: PhoneService, InternetService, OnlineSecurity, etc.
        • Survey responses: Often encoded as 0 (No), 1 (Yes), or 1-5 for ratings
          • Target variable: Churn (has the customer left?)
# Example 3: Count responses for a Likert-scale question (e.g., OnlineSecurity)
likert_counts = df['OnlineSecurity'].value_counts(dropna=False)
print(likert_counts)
OnlineSecurity
No                     3498
Yes                    2019
No internet service    1526
Name: count, dtype: int64
# Example 4: Calculate the percentage distribution for a Likert-scale question
likert_percent = df['OnlineSecurity'].value_counts(normalize=True, dropna=False) * 100
print(likert_percent.round(2))
OnlineSecurity
No                     49.67
Yes                    28.67
No internet service    21.67
Name: proportion, dtype: float64
# Example 5: Visualize Likert-scale ratings (bar plot)
import matplotlib.pyplot as plt
likert_percent.plot(kind='bar', color='skyblue')
plt.title('Online Security - Survey Responses (%)')
plt.ylabel('Percentage')
plt.xlabel('Response')
plt.show()
No description has been provided for this image
# Example 6: Calculate mean satisfaction rating for a numeric scale (e.g., MonthlyCharges)
mean_charge = df['MonthlyCharges'].mean()
print(f'Average Monthly Charges: ${mean_charge:.2f}')
Average Monthly Charges: $64.76
# Example 7: Segment Likert-scale ratings by customer demographics (e.g., gender)
grouped = df.groupby('gender')['OnlineSecurity'].value_counts(normalize=True).unstack().fillna(0) * 100
print(grouped.round(1))
OnlineSecurity    No  No internet service   Yes
gender                                         
Female          49.1                 21.4  29.4
Male            50.2                 21.9  27.9
# Example 8: Cross-tabulate two categorical variables (OnlineSecurity and Churn)
ct = pd.crosstab(df['OnlineSecurity'], df['Churn'], normalize='columns') * 100
print(ct.round(1))
Churn                  No   Yes
OnlineSecurity                 
No                   39.4  78.2
No internet service  27.3   6.0
Yes                  33.3  15.8
# Example 9: Create a satisfaction index from several Likert-scale columns
cols = ['OnlineSecurity', 'TechSupport', 'DeviceProtection']
df['satisfaction_index'] = df[cols].applymap(lambda x: 1 if x == 'Yes' else 0).mean(axis=1)
print(df[['satisfaction_index']].head())
   satisfaction_index
0            0.000000
1            0.666667
2            0.333333
3            1.000000
4            0.000000
# Example 10: Trend analysis for satisfaction index (by tenure in years)
df['tenure_years'] = (df['tenure'] // 12).astype(int)
trend = df.groupby('tenure_years')['satisfaction_index'].mean()
print(trend)
tenure_years
0    0.130014
1    0.224451
2    0.297184
3    0.354278
4    0.389837
5    0.511151
6    0.662063
Name: satisfaction_index, dtype: float64
# Example 11: Advanced - Filter for survey non-response / missing data
missing = df[cols].isnull().sum()
print('Missing responses in key satisfaction columns:')
print(missing)
Missing responses in key satisfaction columns:
OnlineSecurity      0
TechSupport         0
DeviceProtection    0
dtype: int64
# Example 12: Advanced - Handle missing values by filling with 'No Response'
df_filled = df.copy()
df_filled[cols] = df_filled[cols].fillna('No Response')
print(df_filled[cols].head())
  OnlineSecurity TechSupport DeviceProtection
0             No          No               No
1            Yes          No              Yes
2            Yes          No               No
3            Yes         Yes              Yes
4             No          No               No
# Example 13: Advanced - Misinterpreted Likert scales: reverse coding
df['Feedback_Quality'] = np.random.choice([1,2,3,4,5], size=len(df))
df['Quality_reversed'] = 6 - df['Feedback_Quality']
print(df[['Feedback_Quality','Quality_reversed']].head())
   Feedback_Quality  Quality_reversed
0                 3                 3
1                 4                 2
2                 3                 3
3                 4                 2
4                 1                 5
# Example 14: Error Handling - Wrong groupby column name
try:
    df.groupby('gendre')['OnlineSecurity'].mean()
except Exception as e:
    print('Error:', e)
Error: 'gendre'
# Example 15: Error Handling - Interpreting NPS scores as Likert scale
nps_df = pd.DataFrame({'NPS_Score': np.random.randint(0, 11, 100)})
print(nps_df['NPS_Score'].describe())
print('Mean NPS: Not directly interpretable as Net Promoter Score, use the correct classification!')
count    100.000000
mean       5.430000
std        3.188537
min        0.000000
25%        2.750000
50%        6.000000
75%        8.000000
max       10.000000
Name: NPS_Score, dtype: float64
Mean NPS: Not directly interpretable as Net Promoter Score, use the correct classification!
# Example 16: Best Practice - Calculate NPS breakdown
nps_df['Type'] = np.where(nps_df['NPS_Score'] >= 9, 'Promoter',
                      np.where(nps_df['NPS_Score'] <= 6, 'Detractor', 'Passive'))
nps_counts = nps_df['Type'].value_counts(normalize=True) * 100
print('NPS Classification (%):')
print(nps_counts.round(1))
NPS Classification (%):
Type
Detractor    56.0
Passive      23.0
Promoter     21.0
Name: proportion, dtype: float64

Segmentation, cross-tabs, and trend analysis#

  • Key analytics patterns for market research:
    • Segmentation: split responses by demographic or behavioral groups
      • Cross-tabulation: relate two categorical variables to show associations
        • Index or score construction: combine multiple responses into one strong metric
          • Trend analysis: track changes across time or tenure

          • Applying these patterns helps surface business-critical insights.

# Example 17: Best Practice - Segmentation by SeniorCitizen and satisfaction
seg = df.groupby('SeniorCitizen')['satisfaction_index'].mean()
print('Average satisfaction index by senior citizen status:')
print(seg.round(3))
Average satisfaction index by senior citizen status:
SeniorCitizen
0    0.309
1    0.294
Name: satisfaction_index, dtype: float64
# Example 18: Best Practice - Multi-way cross-tab (Churn by InternetService and satisfaction)
df['satisfaction_bin'] = pd.cut(df['satisfaction_index'], bins=[-0.1,0.33,0.66,1], labels=['Low','Medium','High'])
crosstab = pd.crosstab([df['InternetService'], df['satisfaction_bin']], df['Churn'], normalize='columns') * 100
print(crosstab.round(1))
Churn                               No   Yes
InternetService satisfaction_bin            
DSL             Low                6.9  11.7
                Medium            10.7   8.6
                High              20.3   4.3
Fiber optic     Low                9.0  37.2
                Medium            12.5  22.7
                High              13.3   9.4
No              Low               27.3   6.0
# Example 19: Trend chart - satisfaction index by tenure
trend.plot(marker='o')
plt.xlabel('Tenure (years)')
plt.ylabel('Avg Satisfaction Index')
plt.title('Satisfaction Trend by Customer Tenure')
plt.show()
No description has been provided for this image
# Example 20: Tiny end-to-end market research workflow summary
def survey_analysis_workflow(df):
    # Step 1: Check missing values
    missing = df[cols].isnull().sum().sum()
    # Step 2: Compute satisfaction index
    satisfaction = df[cols].applymap(lambda x: 1 if x=='Yes' else 0).mean(axis=1)
    # Step 3: Split by churn
    churned = satisfaction[df['Churn']=='Yes'].mean()
    not_churned = satisfaction[df['Churn']=='No'].mean()
    print(f'Total missing responses: {missing}')
    print(f'Churned Customers Avg Satisfaction: {churned:.3f}')
    print(f'Non-Churned Customers Avg Satisfaction: {not_churned:.3f}')
    if churned < not_churned:
        print('Insight: Lower satisfaction is linked to higher churn. Prioritize improvements!')
    else:
        print('No strong satisfaction-churn link found.')

survey_analysis_workflow(df)
Total missing responses: 0
Churned Customers Avg Satisfaction: 0.205
Non-Churned Customers Avg Satisfaction: 0.344
Insight: Lower satisfaction is linked to higher churn. Prioritize improvements!

Summary and what to try next#

  • You have learned to:
    • Analyze and visualize Likert scale and rating survey data
      • Handle missing values and avoid common pitfalls
        • Segment, cross-tabulate, and construct useful analytics indices
          • Connect survey insights to business recommendations

          • For more hands-on survey analytics, review video case studies on YouTube!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.