Mathew K Analytics

Lesson 43 · Market Research Analytics in Python

Basic Sentiment Analysis for Survey Feedback Using Python

In this lesson, we will learn how to analyze customer survey feedback using sentiment analysis. Understanding customer sentiment helps businesses improve…

⬇ Download notebookOpen in Colab ↗

📓 Full notebook

Download .ipynb

Basic Sentiment Analysis for Survey Feedback#

  • In this lesson, we will learn how to analyze customer survey feedback using sentiment analysis.
  • Understanding customer sentiment helps businesses improve services, products, and customer experiences.
  • We will use real-world datasets to identify trends in customer opinions through survey responses and open-ended feedback.
  • By the end, you will be able to turn unstructured survey text into actionable insights for market research and customer analytics.
import pandas as pd
import numpy as np
import openml
import warnings
warnings.filterwarnings('ignore')

Core Concepts: Survey Data and Sentiment Analysis#

  • Customer surveys often include both closed-ended scores (like NPS or Likert scales) and open-ended feedback.
  • Each row in such data represents one customer's responses, with columns for demographic and survey details.
  • Feedback text may include typos, neutral phrases, or even sarcasm, which can be hard for simple analysis.
  • Beginners often miss cleaning, filtering, or handling missing feedback before analyzing sentiment.
  • It is important to separate objective ratings from subjective opinions for accurate business insights.
df = pd.DataFrame({'CustomerID':[1,2,3,4,5], 'Feedback':['Great service and friendly staff','Delivery was slow and packaging was poor','Excellent quality, will buy again','Customer support needs improvement','Good value for money']})
print(df.shape)
print(df.head(3))
(5, 2)
   CustomerID                                  Feedback
0           1          Great service and friendly staff
1           2  Delivery was slow and packaging was poor
2           3         Excellent quality, will buy again

Beginner Example 1: Viewing All Customer Feedback#

  • Let us preview all open-ended customer feedback comments.
  • Observing raw comments is important for spotting common themes or issues.
for idx, row in df.iterrows():
    print(f"Customer {row['CustomerID']}: {row['Feedback']}")
Customer 1: Great service and friendly staff
Customer 2: Delivery was slow and packaging was poor
Customer 3: Excellent quality, will buy again
Customer 4: Customer support needs improvement
Customer 5: Good value for money

Beginner Example 2: Adding a Simple Sentiment Label#

  • We will add a simple sentiment label using keywords.
  • This quick approach helps classify feedback as Positive, Negative, or Neutral.
  • It is common in business settings as a first pass for large text datasets.
def simple_sentiment(text):
    text = text.lower()
    if any(word in text for word in ['great', 'excellent', 'good', 'friendly']):
        return 'Positive'
    elif any(word in text for word in ['slow', 'poor', 'needs improvement']):
        return 'Negative'
    else:
        return 'Neutral'

df['Sentiment'] = df['Feedback'].apply(simple_sentiment)
print(df[['Feedback','Sentiment']])
                                   Feedback Sentiment
0          Great service and friendly staff  Positive
1  Delivery was slow and packaging was poor  Negative
2         Excellent quality, will buy again  Positive
3        Customer support needs improvement  Negative
4                      Good value for money  Positive

Beginner Example 3: Visualizing Sentiment Distribution#

  • Visualizing the distribution of sentiment labels helps us see the overall mood in customer responses.
  • Pie charts or bar graphs are common for such comparisons.
import matplotlib.pyplot as plt
sent_counts = df['Sentiment'].value_counts()
sent_counts.plot(kind='bar', color=['green','red','gray'])
plt.title('Sentiment Distribution in Customer Feedback')
plt.ylabel('Number of Comments')
plt.show()
No description has been provided for this image

Intermediate Example 1: Analyzing Larger Survey Data (NPS)#

  • Net Promoter Score (NPS) surveys ask how likely a customer is to recommend the company.
  • Scores range from 0-10 and can be linked with open-ended feedback for deeper analysis.
np.random.seed(42)
nps_df = pd.DataFrame({
    'CustomerID': range(6,16),
    'NPS_Score': np.random.randint(0, 11, 10),
    'Feedback': [
        'Loved the app and interface',
        'Not satisfied with delivery times',
        'Staff was super helpful',
        'The product broke quickly',
        'Amazing experience',
        'Average service overall',
        'Very fast and simple checkout',
        'Difficult to reach support',
        'Fantastic offers and promotions',
        'Unclear billing statements'
    ]
})
print(nps_df.head())
   CustomerID  NPS_Score                           Feedback
0           6          6        Loved the app and interface
1           7          3  Not satisfied with delivery times
2           8         10            Staff was super helpful
3           9          7          The product broke quickly
4          10          4                 Amazing experience

Intermediate Example 2: Linking NPS Score and Text Sentiment#

  • By combining text sentiment and NPS scores, we can spot mismatches: e.g., negative text with high scores.
  • This helps us catch cases where scores and text do not match business expectations.
nps_df['Sentiment'] = nps_df['Feedback'].apply(simple_sentiment)
mismatch = nps_df[((nps_df['NPS_Score'] > 7) & (nps_df['Sentiment'] == 'Negative')) | ((nps_df['NPS_Score'] < 4) & (nps_df['Sentiment'] == 'Positive'))]
print('Comments where score and sentiment do not match:')
print(mismatch[['NPS_Score','Feedback','Sentiment']])
Comments where score and sentiment do not match:
Empty DataFrame
Columns: [NPS_Score, Feedback, Sentiment]
Index: []

Intermediate Example 3: Counting Promoters, Passives, and Detractors#

  • In NPS, 0-6 are Detractors, 7-8 are Passives, 9-10 are Promoters.
  • Segmenting customers helps businesses target improvements to the right group.
def nps_category(score):
    if score <= 6: return 'Detractor'
    elif score <= 8: return 'Passive'
    else: return 'Promoter'
nps_df['NPS_Type'] = nps_df['NPS_Score'].apply(nps_category)
print(nps_df['NPS_Type'].value_counts())
NPS_Type
Detractor    6
Promoter     3
Passive      1
Name: count, dtype: int64

Advanced Example 1: Sentiment Analysis on Larger Survey Feedback#

  • Let us load a real customer satisfaction dataset from OpenML, which includes demographics and service ratings.
  • We will analyze open-ended feedback to guide business decisions.
dataset = openml.datasets.get_dataset(42178)
sf_df, _, _, _ = dataset.get_data(dataset_format='dataframe')
print(sf_df.shape)
print(sf_df.head(3))
(7043, 20)
   gender  SeniorCitizen Partner Dependents  tenure PhoneService  \
0  Female              0     Yes         No       1           No   
1    Male              0      No         No      34          Yes   
2    Male              0      No         No       2          Yes   

      MultipleLines InternetService OnlineSecurity OnlineBackup  \
0  No phone service             DSL             No          Yes   
1                No             DSL            Yes           No   
2                No             DSL            Yes          Yes   

  DeviceProtection TechSupport StreamingTV StreamingMovies        Contract  \
0               No          No          No              No  Month-to-month   
1              Yes          No          No              No        One year   
2               No          No          No              No  Month-to-month   

  PaperlessBilling     PaymentMethod  MonthlyCharges TotalCharges Churn  
0              Yes  Electronic check           29.85        29.85    No  
1               No      Mailed check           56.95       1889.5    No  
2              Yes      Mailed check           53.85       108.15   Yes  

Advanced Example 2: Handling and Analyzing Missing Feedback#

  • Not all customers provide textual feedback. Missing values can mislead your market analysis.
  • We must detect and handle missing comments before analyzing sentiment distributions.
num_missing = df['Feedback'].isnull().sum()
print(f"Missing feedback entries: {num_missing}")
df['Feedback_filled'] = df['Feedback'].fillna('No Feedback Provided')
print(df[['Feedback','Feedback_filled']].head(3))
Missing feedback entries: 0
                                   Feedback  \
0          Great service and friendly staff   
1  Delivery was slow and packaging was poor   
2         Excellent quality, will buy again   

                            Feedback_filled  
0          Great service and friendly staff  
1  Delivery was slow and packaging was poor  
2         Excellent quality, will buy again  

Error Handling: Incorrect Groupings or Aggregations#

  • Mistakes can occur if you group by the wrong column, leading to misleading market insights.
  • Always double-check groupings when calculating sentiment across customer segments.
# Wrong: grouping by CustomerID (should use Sentiment)
bad_group = df.groupby('CustomerID')['Sentiment'].count()
print('Wrong grouping:')
print(bad_group.head(3))

# Correct: grouping by Sentiment
good_group = df.groupby('Sentiment')['CustomerID'].count()
print('\nCorrect grouping:')
print(good_group)
Wrong grouping:
CustomerID
1    1
2    1
3    1
Name: Sentiment, dtype: int64

Correct grouping:
Sentiment
Negative    2
Positive    3
Name: CustomerID, dtype: int64

Error Handling: Misinterpreting Scales#

  • Treat all response options carefully. For example, in NPS, the neutral point is not 5, it is 7-8.
  • Review official scoring to avoid wrong business conclusions.
nps_df['Guessed_Neutral'] = nps_df['NPS_Score'] == 5
real_neutral = nps_df['NPS_Type'] == 'Passive'
print(f"Guessed neutral (score=5): {nps_df['Guessed_Neutral'].sum()}")
print(f"True passives (score 7-8): {real_neutral.sum()}")
Guessed neutral (score=5): 0
True passives (score 7-8): 1

Best Practices: Market Research Analytics Patterns#

  • Segmenting by demographics or region uncovers hidden sentiment trends among customer subgroups.
  • Cross-tabulate sentiment versus churn, region, or NPS type for actionable insights.
  • Constructing composite indexes (like overall satisfaction or loyalty scores) can clarify trends.
  • Analyzing sentiment or NPS changes over time reveals whether business actions are working.
np.random.seed(42)
nps_survey = pd.DataFrame({'CustomerID': range(1, 21),
                         'Age': np.random.randint(18, 70, 20),
                         'Region': np.random.choice(['North', 'South', 'East', 'West'], 20),
                         'NPS_Score': np.random.randint(0, 11, 20),
                         'Feedback': np.random.choice(df['Feedback'], 20)})
nps_survey['Sentiment'] = nps_survey['Feedback'].apply(simple_sentiment)
grouped = nps_survey.groupby('Region')['Sentiment'].value_counts().unstack().fillna(0)
print('Sentiment by Region:')
print(grouped)
Sentiment by Region:
Sentiment  Negative  Positive
Region                       
East            3.0       0.0
North           3.0       2.0
South           6.0       0.0
West            4.0       2.0
sentiment_by_score = nps_survey.groupby('Sentiment')['NPS_Score'].mean()
print('Average NPS Score by Sentiment:')
print(sentiment_by_score)
Average NPS Score by Sentiment:
Sentiment
Negative    4.9375
Positive    4.0000
Name: NPS_Score, dtype: float64
trend_df = nps_survey.copy()
trend_df['SurveyMonth'] = np.random.choice(['2023-03','2023-04','2023-05'], len(trend_df))
monthly_sent = trend_df.groupby('SurveyMonth')['Sentiment'].value_counts().unstack().fillna(0)
print('Monthly sentiment counts:')
print(monthly_sent)
monthly_sent.plot(kind='bar', stacked=True)
plt.title('Monthly Sentiment Trends in Feedback')
plt.ylabel('Number of Comments')
plt.show()
Monthly sentiment counts:
Sentiment    Negative  Positive
SurveyMonth                    
2023-03             8         2
2023-04             4         1
2023-05             4         1
No description has been provided for this image

End-to-End Example: Generating a Market Research Recommendation#

  • Let us put it all together: analyze survey feedback and produce a recommendation for the business.
  • We will summarize sentiment, segment results, and suggest an action based on the data.
summary = nps_survey['Sentiment'].value_counts(normalize=True) * 100
rec = ''
if summary.get('Negative', 0) > 30:
    rec = 'Focus on service improvements and proactive support.'
elif summary.get('Positive', 0) > 60:
    rec = 'Leverage positive sentiment in marketing campaigns.'
else:
    rec = 'Monitor feedback and maintain current service levels.'
print('Sentiment Summary (%):')
print(summary.round(1))
print('Business Recommendation:')
print(rec)
Sentiment Summary (%):
Sentiment
Negative    80.0
Positive    20.0
Name: proportion, dtype: float64
Business Recommendation:
Focus on service improvements and proactive support.
 

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.