Lesson 44 · Market Research Analytics in Python
Identifying Common Themes in Customer Comments with Python Analytics
In market research, companies collect open-ended customer feedback to learn what people feel and think. If we can identify common themes, we make better…
- CourseMarket Research Analytics in Python
- Lesson44 of 56
- Video26 min
- FormatJupyter notebook · 26 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbIdentifying Common Themes in Customer Comments#
- In market research, companies collect open-ended customer feedback to learn what people feel and think.
- If we can identify common themes, we make better business decisions, improve services, and reduce churn.
- Today we will analyze real feedback data step by step and discover insights your team can act on.
- You will learn practical skills to summarize feedback, group similar comments, and create actionable recommendations from raw text.
import pandas as pd
import numpy as np
import warnings
warnings.filterwarnings('ignore')
Understanding customer feedback data#
- Customer feedback usually comes as free-text comments or open-ended survey responses.
- Good research often combines numeric scores (1-10, NPS) with customer written remarks.
- Some datasets have rows for each comment, while others join all feedback into one field.
- Beginners often forget to clean text (remove punctuation, lowercasing) or handle missing feedback.
- Handle duplicate or boilerplate comments carefully to avoid biasing the results.
- Always check the context: is it a complaint, a praise, or a suggestion?
df = pd.DataFrame({
'CustomerID':[1,2,3,4,5],
'Feedback':[
'Great service and friendly staff',
'Delivery was slow and packaging was poor',
'Excellent quality, will buy again',
'Customer support needs improvement',
'Good value for money']})
print(df.shape)
print(df.head(3))
df['Feedback_clean'] = df['Feedback'].str.lower().str.replace(r'[^a-z\s]', '', regex=True)
print(df[['Feedback','Feedback_clean']])
df['Feedback_wordcount'] = df['Feedback_clean'].str.split().apply(len)
print(df[['Feedback','Feedback_wordcount']])
all_words = pd.Series(' '.join(df['Feedback_clean']).split())
print(all_words.value_counts().head(10))
stopwords = ['and','was','for','the','in','is','to','of','a','an']
all_words_no_stop = all_words[~all_words.isin(stopwords)]
print(all_words_no_stop.value_counts().head(10))
import matplotlib.pyplot as plt
word_counts = all_words_no_stop.value_counts().head(5)
plt.bar(word_counts.index, word_counts.values)
plt.ylabel('Count')
plt.title('Top 5 Most Common Feedback Words')
plt.show()
theme_keywords = {
'service':['service','support','staff','help'],
'delivery':['delivery','packaging','slow'],
'quality':['quality','excellent','good'],
'value':['value','money'],
'improvement':['improvement','needs'],
}
def tag_theme(text):
out = []
for theme,words in theme_keywords.items():
if any(w in text for w in words):
out.append(theme)
return ', '.join(out) if out else 'other'
df['Theme'] = df['Feedback_clean'].apply(tag_theme)
print(df[['Feedback','Theme']])
Intermediate: Group and count by theme#
- Now that we have themes assigned, we can count how many comments belong to each theme.
- This tells us which business areas are mentioned most.
theme_counts = df['Theme'].value_counts()
print(theme_counts)
theme_counts.plot(kind='bar',title='Number of Comments per Theme')
plt.ylabel('Number of Comments')
plt.show()
grouped = df.groupby('Theme').agg({'CustomerID':'count','Feedback_wordcount':'mean'})
print(grouped.rename(columns={'CustomerID':'Comment Count','Feedback_wordcount':'Avg Words'}))
nps_df = pd.DataFrame({
'CustomerID': [1,2,3,4,5],
'NPS_Score': [9, 3, 10, 5, 7]
})
result = df.merge(nps_df, on='CustomerID')
print(result[['Feedback','Theme','NPS_Score']])
# Bucket NPS into promoters, passives, detractors
def nps_group(x):
if x >= 9:
return 'Promoter'
elif x >= 7:
return 'Passive'
else:
return 'Detractor'
result['NPS_Group'] = result['NPS_Score'].apply(nps_group)
summary = result.groupby(['Theme','NPS_Group']).size().unstack().fillna(0)
print(summary)
Advanced: Theme extraction on a larger real dataset#
- To practice on more complex and realistic data, let us load a truly open-ended customer feedback dataset.
- We will clean the text, extract keywords, and use more automated grouping methods.
real_df = pd.DataFrame({
'CustomerID':[101,102,103,104,105,106,107,108],
'Feedback':[
'The checkout process was confusing and I could not apply my coupon',
'Shipping was extremely fast and everything arrived in perfect condition',
'I had trouble finding the sizing guide on mobile',
'Would recommend. Good quality and price!',
'My issue took too long to resolve with support',
'The product is not as pictured',
'Very pleasant shopping experience',
'Got my refund right away, thank you!'
]
})
real_df['Feedback_clean'] = real_df['Feedback'].str.lower().str.replace(r'[^a-z\s]', '', regex=True)
print(real_df.head(3))
real_all_words = pd.Series(' '.join(real_df['Feedback_clean']).split())
real_stopwords = set(['and','the','i','in','my','to','with','was','had','is','on','for','not','as','could','a','of','it'])
words_no_stop = real_all_words[~real_all_words.isin(real_stopwords)]
print(words_no_stop.value_counts().head(10))
from collections import Counter
themes_advanced = {
'shipping':['shipping','fast','arrived'],
'support':['support','issue','refund','resolve','thank'],
'website':['checkout','process','coupon','mobile','finding','guide','confusing'],
'quality':['perfect','good','quality','pictured','product'],
'experience':['pleasant','recommend','shopping','experience']
}
def extract_theme_adv(text):
found = []
for t,words in themes_advanced.items():
if any(w in text for w in words):
found.append(t)
return ', '.join(found) if found else 'other'
real_df['Theme'] = real_df['Feedback_clean'].apply(extract_theme_adv)
print(real_df[['Feedback','Theme']])
theme_dist = real_df['Theme'].str.get_dummies(sep=', ').sum(axis=0).sort_values(ascending=False)
print(theme_dist)
theme_dist.plot(kind='bar',title='Theme Mentions in Customer Comments',color='steelblue')
plt.ylabel('Mentions')
plt.show()
Error handling and debugging (missing feedback)#
- Real-world survey data often has missing written comments.
- We must detect and handle missing or empty feedback to avoid bias or errors.
test_df = pd.DataFrame({
'CustomerID':[11,12,13],
'Feedback':['Great service',np.nan,'']
})
missing = test_df['Feedback'].isnull() | (test_df['Feedback'].str.strip() == '')
print('Missing feedback entries:')
print(test_df[missing])
sample_df = pd.DataFrame({'Theme':['service','delivery','service','value','service']})
wrong_counts = sample_df.groupby('Theme').count()
print('Incorrect grouping result:')
print(wrong_counts)
right_counts = sample_df.groupby('Theme').size()
print('Correct theme counts:')
print(right_counts)
likert_scores = pd.Series(['Strongly agree', 'Agree', 'Neutral', 'Disagree', 'Strongly agree'])
score_map = {
'Strongly agree':5,
'Agree':4,
'Neutral':3,
'Disagree':2,
'Strongly disagree':1
}
num_scores = likert_scores.map(score_map)
print(num_scores)
Best practices for analyzing open feedback#
- Segment comments by customer type or sentiment, not just topic.
- Cross-tabulate themes with NPS groups, age, or purchase size.
- Create composite indices or theme scores if management wants summary KPIs.
- Track trends in complaint or praise themes over time for early warnings.
# Example: Cross-tabulate theme by NPS group from earlier
print(result.groupby(['Theme','NPS_Group']).size().unstack().fillna(0))
# Build a theme score (1 point for each theme per comment)
result['Theme_Score'] = result['Theme'].apply(lambda x: 0 if x=='other' else len(x.split(',')))
print(result[['Feedback','Theme','Theme_Score']])
# Example: Trend analysis (simulate monthly feedback counts by theme)
trend_df = pd.DataFrame({
'Month':['2023-01','2023-01','2023-02','2023-02','2023-02','2023-03','2023-03'],
'Theme':['service','delivery','service','delivery','value','value','service']
})
pivot = trend_df.pivot_table(index='Month',columns='Theme',values='Theme',aggfunc='count').fillna(0)
print(pivot)
End-to-end market research insight: What can we recommend?#
- Collect raw feedback data and assign business themes using keywords.
- Link themes to customer advocacy or satisfaction groups.
- Compute the most common complaint and praise topics.
- Recommend actionable changes where themes show high volume or negative sentiment.
- Summarize findings for business teams in charts and tables.
# Workflow: From raw data to recommended business actions
final_theme_counts = real_df['Theme'].str.get_dummies(sep=', ').sum(axis=0)
main_problem = final_theme_counts.idxmax()
print('The most common feedback theme is:', main_problem)
if main_problem in ['shipping','support','website']:
print('Business should review processes in', main_problem, 'immediately!')
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



