Lesson 43 · Market Research Analytics in Python
Basic Sentiment Analysis for Survey Feedback Using Python
In this lesson, we will learn how to analyze customer survey feedback using sentiment analysis. Understanding customer sentiment helps businesses improve…
- CourseMarket Research Analytics in Python
- Lesson43 of 56
- Video19 min
- FormatJupyter notebook · 17 code cells
What you'll learn
- Core Concepts: Survey Data and Sentiment Analysis
- Beginner Example 1: Viewing All Customer Feedback
- Beginner Example 2: Adding a Simple Sentiment Label
- Beginner Example 3: Visualizing Sentiment Distribution
- Intermediate Example 1: Analyzing Larger Survey Data (NPS)
- Intermediate Example 2: Linking NPS Score and Text Sentiment
- Intermediate Example 3: Counting Promoters, Passives, and Detractors
- Advanced Example 1: Sentiment Analysis on Larger Survey Feedback
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbBasic Sentiment Analysis for Survey Feedback#
- In this lesson, we will learn how to analyze customer survey feedback using sentiment analysis.
- Understanding customer sentiment helps businesses improve services, products, and customer experiences.
- We will use real-world datasets to identify trends in customer opinions through survey responses and open-ended feedback.
- By the end, you will be able to turn unstructured survey text into actionable insights for market research and customer analytics.
import pandas as pd
import numpy as np
import openml
import warnings
warnings.filterwarnings('ignore')
Core Concepts: Survey Data and Sentiment Analysis#
- Customer surveys often include both closed-ended scores (like NPS or Likert scales) and open-ended feedback.
- Each row in such data represents one customer's responses, with columns for demographic and survey details.
- Feedback text may include typos, neutral phrases, or even sarcasm, which can be hard for simple analysis.
- Beginners often miss cleaning, filtering, or handling missing feedback before analyzing sentiment.
- It is important to separate objective ratings from subjective opinions for accurate business insights.
df = pd.DataFrame({'CustomerID':[1,2,3,4,5], 'Feedback':['Great service and friendly staff','Delivery was slow and packaging was poor','Excellent quality, will buy again','Customer support needs improvement','Good value for money']})
print(df.shape)
print(df.head(3))
Beginner Example 1: Viewing All Customer Feedback#
- Let us preview all open-ended customer feedback comments.
- Observing raw comments is important for spotting common themes or issues.
for idx, row in df.iterrows():
print(f"Customer {row['CustomerID']}: {row['Feedback']}")
Beginner Example 2: Adding a Simple Sentiment Label#
- We will add a simple sentiment label using keywords.
- This quick approach helps classify feedback as Positive, Negative, or Neutral.
- It is common in business settings as a first pass for large text datasets.
def simple_sentiment(text):
text = text.lower()
if any(word in text for word in ['great', 'excellent', 'good', 'friendly']):
return 'Positive'
elif any(word in text for word in ['slow', 'poor', 'needs improvement']):
return 'Negative'
else:
return 'Neutral'
df['Sentiment'] = df['Feedback'].apply(simple_sentiment)
print(df[['Feedback','Sentiment']])
Beginner Example 3: Visualizing Sentiment Distribution#
- Visualizing the distribution of sentiment labels helps us see the overall mood in customer responses.
- Pie charts or bar graphs are common for such comparisons.
import matplotlib.pyplot as plt
sent_counts = df['Sentiment'].value_counts()
sent_counts.plot(kind='bar', color=['green','red','gray'])
plt.title('Sentiment Distribution in Customer Feedback')
plt.ylabel('Number of Comments')
plt.show()
Intermediate Example 1: Analyzing Larger Survey Data (NPS)#
- Net Promoter Score (NPS) surveys ask how likely a customer is to recommend the company.
- Scores range from 0-10 and can be linked with open-ended feedback for deeper analysis.
np.random.seed(42)
nps_df = pd.DataFrame({
'CustomerID': range(6,16),
'NPS_Score': np.random.randint(0, 11, 10),
'Feedback': [
'Loved the app and interface',
'Not satisfied with delivery times',
'Staff was super helpful',
'The product broke quickly',
'Amazing experience',
'Average service overall',
'Very fast and simple checkout',
'Difficult to reach support',
'Fantastic offers and promotions',
'Unclear billing statements'
]
})
print(nps_df.head())
Intermediate Example 2: Linking NPS Score and Text Sentiment#
- By combining text sentiment and NPS scores, we can spot mismatches: e.g., negative text with high scores.
- This helps us catch cases where scores and text do not match business expectations.
nps_df['Sentiment'] = nps_df['Feedback'].apply(simple_sentiment)
mismatch = nps_df[((nps_df['NPS_Score'] > 7) & (nps_df['Sentiment'] == 'Negative')) | ((nps_df['NPS_Score'] < 4) & (nps_df['Sentiment'] == 'Positive'))]
print('Comments where score and sentiment do not match:')
print(mismatch[['NPS_Score','Feedback','Sentiment']])
Intermediate Example 3: Counting Promoters, Passives, and Detractors#
- In NPS, 0-6 are Detractors, 7-8 are Passives, 9-10 are Promoters.
- Segmenting customers helps businesses target improvements to the right group.
def nps_category(score):
if score <= 6: return 'Detractor'
elif score <= 8: return 'Passive'
else: return 'Promoter'
nps_df['NPS_Type'] = nps_df['NPS_Score'].apply(nps_category)
print(nps_df['NPS_Type'].value_counts())
Advanced Example 1: Sentiment Analysis on Larger Survey Feedback#
- Let us load a real customer satisfaction dataset from OpenML, which includes demographics and service ratings.
- We will analyze open-ended feedback to guide business decisions.
dataset = openml.datasets.get_dataset(42178)
sf_df, _, _, _ = dataset.get_data(dataset_format='dataframe')
print(sf_df.shape)
print(sf_df.head(3))
Advanced Example 2: Handling and Analyzing Missing Feedback#
- Not all customers provide textual feedback. Missing values can mislead your market analysis.
- We must detect and handle missing comments before analyzing sentiment distributions.
num_missing = df['Feedback'].isnull().sum()
print(f"Missing feedback entries: {num_missing}")
df['Feedback_filled'] = df['Feedback'].fillna('No Feedback Provided')
print(df[['Feedback','Feedback_filled']].head(3))
Error Handling: Incorrect Groupings or Aggregations#
- Mistakes can occur if you group by the wrong column, leading to misleading market insights.
- Always double-check groupings when calculating sentiment across customer segments.
# Wrong: grouping by CustomerID (should use Sentiment)
bad_group = df.groupby('CustomerID')['Sentiment'].count()
print('Wrong grouping:')
print(bad_group.head(3))
# Correct: grouping by Sentiment
good_group = df.groupby('Sentiment')['CustomerID'].count()
print('\nCorrect grouping:')
print(good_group)
Error Handling: Misinterpreting Scales#
- Treat all response options carefully. For example, in NPS, the neutral point is not 5, it is 7-8.
- Review official scoring to avoid wrong business conclusions.
nps_df['Guessed_Neutral'] = nps_df['NPS_Score'] == 5
real_neutral = nps_df['NPS_Type'] == 'Passive'
print(f"Guessed neutral (score=5): {nps_df['Guessed_Neutral'].sum()}")
print(f"True passives (score 7-8): {real_neutral.sum()}")
Best Practices: Market Research Analytics Patterns#
- Segmenting by demographics or region uncovers hidden sentiment trends among customer subgroups.
- Cross-tabulate sentiment versus churn, region, or NPS type for actionable insights.
- Constructing composite indexes (like overall satisfaction or loyalty scores) can clarify trends.
- Analyzing sentiment or NPS changes over time reveals whether business actions are working.
np.random.seed(42)
nps_survey = pd.DataFrame({'CustomerID': range(1, 21),
'Age': np.random.randint(18, 70, 20),
'Region': np.random.choice(['North', 'South', 'East', 'West'], 20),
'NPS_Score': np.random.randint(0, 11, 20),
'Feedback': np.random.choice(df['Feedback'], 20)})
nps_survey['Sentiment'] = nps_survey['Feedback'].apply(simple_sentiment)
grouped = nps_survey.groupby('Region')['Sentiment'].value_counts().unstack().fillna(0)
print('Sentiment by Region:')
print(grouped)
sentiment_by_score = nps_survey.groupby('Sentiment')['NPS_Score'].mean()
print('Average NPS Score by Sentiment:')
print(sentiment_by_score)
trend_df = nps_survey.copy()
trend_df['SurveyMonth'] = np.random.choice(['2023-03','2023-04','2023-05'], len(trend_df))
monthly_sent = trend_df.groupby('SurveyMonth')['Sentiment'].value_counts().unstack().fillna(0)
print('Monthly sentiment counts:')
print(monthly_sent)
monthly_sent.plot(kind='bar', stacked=True)
plt.title('Monthly Sentiment Trends in Feedback')
plt.ylabel('Number of Comments')
plt.show()
End-to-End Example: Generating a Market Research Recommendation#
- Let us put it all together: analyze survey feedback and produce a recommendation for the business.
- We will summarize sentiment, segment results, and suggest an action based on the data.
summary = nps_survey['Sentiment'].value_counts(normalize=True) * 100
rec = ''
if summary.get('Negative', 0) > 30:
rec = 'Focus on service improvements and proactive support.'
elif summary.get('Positive', 0) > 60:
rec = 'Leverage positive sentiment in marketing campaigns.'
else:
rec = 'Monitor feedback and maintain current service levels.'
print('Sentiment Summary (%):')
print(summary.round(1))
print('Business Recommendation:')
print(rec)
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



