Mathew K Analytics

Lesson 44 · Market Research Analytics in Python

Identifying Common Themes in Customer Comments with Python Analytics

In market research, companies collect open-ended customer feedback to learn what people feel and think. If we can identify common themes, we make better…

⬇ Download notebookOpen in Colab ↗

📓 Full notebook

Download .ipynb

Identifying Common Themes in Customer Comments#

  • In market research, companies collect open-ended customer feedback to learn what people feel and think.
  • If we can identify common themes, we make better business decisions, improve services, and reduce churn.
  • Today we will analyze real feedback data step by step and discover insights your team can act on.
  • You will learn practical skills to summarize feedback, group similar comments, and create actionable recommendations from raw text.
import pandas as pd
import numpy as np
import warnings
warnings.filterwarnings('ignore')

Understanding customer feedback data#

  • Customer feedback usually comes as free-text comments or open-ended survey responses.
  • Good research often combines numeric scores (1-10, NPS) with customer written remarks.
  • Some datasets have rows for each comment, while others join all feedback into one field.
  • Beginners often forget to clean text (remove punctuation, lowercasing) or handle missing feedback.
  • Handle duplicate or boilerplate comments carefully to avoid biasing the results.
  • Always check the context: is it a complaint, a praise, or a suggestion?
df = pd.DataFrame({
    'CustomerID':[1,2,3,4,5],
    'Feedback':[
        'Great service and friendly staff',
        'Delivery was slow and packaging was poor',
        'Excellent quality, will buy again',
        'Customer support needs improvement',
        'Good value for money']})
print(df.shape)
print(df.head(3))
(5, 2)
   CustomerID                                  Feedback
0           1          Great service and friendly staff
1           2  Delivery was slow and packaging was poor
2           3         Excellent quality, will buy again
df['Feedback_clean'] = df['Feedback'].str.lower().str.replace(r'[^a-z\s]', '', regex=True)
print(df[['Feedback','Feedback_clean']])
                                   Feedback  \
0          Great service and friendly staff   
1  Delivery was slow and packaging was poor   
2         Excellent quality, will buy again   
3        Customer support needs improvement   
4                      Good value for money   

                             Feedback_clean  
0          great service and friendly staff  
1  delivery was slow and packaging was poor  
2          excellent quality will buy again  
3        customer support needs improvement  
4                      good value for money  
df['Feedback_wordcount'] = df['Feedback_clean'].str.split().apply(len)
print(df[['Feedback','Feedback_wordcount']])
                                   Feedback  Feedback_wordcount
0          Great service and friendly staff                   5
1  Delivery was slow and packaging was poor                   7
2         Excellent quality, will buy again                   5
3        Customer support needs improvement                   4
4                      Good value for money                   4
all_words = pd.Series(' '.join(df['Feedback_clean']).split())
print(all_words.value_counts().head(10))
and          2
was          2
great        1
service      1
friendly     1
staff        1
delivery     1
slow         1
packaging    1
poor         1
Name: count, dtype: int64
stopwords = ['and','was','for','the','in','is','to','of','a','an']
all_words_no_stop = all_words[~all_words.isin(stopwords)]
print(all_words_no_stop.value_counts().head(10))
great        1
service      1
friendly     1
staff        1
delivery     1
slow         1
packaging    1
poor         1
excellent    1
quality      1
Name: count, dtype: int64
import matplotlib.pyplot as plt
word_counts = all_words_no_stop.value_counts().head(5)
plt.bar(word_counts.index, word_counts.values)
plt.ylabel('Count')
plt.title('Top 5 Most Common Feedback Words')
plt.show()
No description has been provided for this image
theme_keywords = {
    'service':['service','support','staff','help'],
    'delivery':['delivery','packaging','slow'],
    'quality':['quality','excellent','good'],
    'value':['value','money'],
    'improvement':['improvement','needs'],
}

def tag_theme(text):
    out = []
    for theme,words in theme_keywords.items():
        if any(w in text for w in words):
            out.append(theme)
    return ', '.join(out) if out else 'other'

df['Theme'] = df['Feedback_clean'].apply(tag_theme)
print(df[['Feedback','Theme']])
                                   Feedback                 Theme
0          Great service and friendly staff               service
1  Delivery was slow and packaging was poor              delivery
2         Excellent quality, will buy again               quality
3        Customer support needs improvement  service, improvement
4                      Good value for money        quality, value

Intermediate: Group and count by theme#

  • Now that we have themes assigned, we can count how many comments belong to each theme.
  • This tells us which business areas are mentioned most.
theme_counts = df['Theme'].value_counts()
print(theme_counts)
Theme
service                 1
delivery                1
quality                 1
service, improvement    1
quality, value          1
Name: count, dtype: int64
theme_counts.plot(kind='bar',title='Number of Comments per Theme')
plt.ylabel('Number of Comments')
plt.show()
No description has been provided for this image
grouped = df.groupby('Theme').agg({'CustomerID':'count','Feedback_wordcount':'mean'})
print(grouped.rename(columns={'CustomerID':'Comment Count','Feedback_wordcount':'Avg Words'}))
                      Comment Count  Avg Words
Theme                                         
delivery                          1        7.0
quality                           1        5.0
quality, value                    1        4.0
service                           1        5.0
service, improvement              1        4.0
nps_df = pd.DataFrame({
    'CustomerID': [1,2,3,4,5],
    'NPS_Score': [9, 3, 10, 5, 7]
})
result = df.merge(nps_df, on='CustomerID')
print(result[['Feedback','Theme','NPS_Score']])
                                   Feedback                 Theme  NPS_Score
0          Great service and friendly staff               service          9
1  Delivery was slow and packaging was poor              delivery          3
2         Excellent quality, will buy again               quality         10
3        Customer support needs improvement  service, improvement          5
4                      Good value for money        quality, value          7
# Bucket NPS into promoters, passives, detractors
def nps_group(x):
    if x >= 9:
        return 'Promoter'
    elif x >= 7:
        return 'Passive'
    else:
        return 'Detractor'

result['NPS_Group'] = result['NPS_Score'].apply(nps_group)
summary = result.groupby(['Theme','NPS_Group']).size().unstack().fillna(0)
print(summary)
NPS_Group             Detractor  Passive  Promoter
Theme                                             
delivery                    1.0      0.0       0.0
quality                     0.0      0.0       1.0
quality, value              0.0      1.0       0.0
service                     0.0      0.0       1.0
service, improvement        1.0      0.0       0.0

Advanced: Theme extraction on a larger real dataset#

  • To practice on more complex and realistic data, let us load a truly open-ended customer feedback dataset.
  • We will clean the text, extract keywords, and use more automated grouping methods.
real_df = pd.DataFrame({
    'CustomerID':[101,102,103,104,105,106,107,108],
    'Feedback':[
        'The checkout process was confusing and I could not apply my coupon',
        'Shipping was extremely fast and everything arrived in perfect condition',
        'I had trouble finding the sizing guide on mobile',
        'Would recommend. Good quality and price!',
        'My issue took too long to resolve with support',
        'The product is not as pictured',
        'Very pleasant shopping experience',
        'Got my refund right away, thank you!'
    ]
})
real_df['Feedback_clean'] = real_df['Feedback'].str.lower().str.replace(r'[^a-z\s]', '', regex=True)
print(real_df.head(3))
   CustomerID                                           Feedback  \
0         101  The checkout process was confusing and I could...   
1         102  Shipping was extremely fast and everything arr...   
2         103   I had trouble finding the sizing guide on mobile   

                                      Feedback_clean  
0  the checkout process was confusing and i could...  
1  shipping was extremely fast and everything arr...  
2   i had trouble finding the sizing guide on mobile  
real_all_words = pd.Series(' '.join(real_df['Feedback_clean']).split())
real_stopwords = set(['and','the','i','in','my','to','with','was','had','is','on','for','not','as','could','a','of','it'])
words_no_stop = real_all_words[~real_all_words.isin(real_stopwords)]
print(words_no_stop.value_counts().head(10))
checkout      1
process       1
confusing     1
apply         1
coupon        1
shipping      1
extremely     1
fast          1
everything    1
arrived       1
Name: count, dtype: int64
from collections import Counter
themes_advanced = {
    'shipping':['shipping','fast','arrived'],
    'support':['support','issue','refund','resolve','thank'],
    'website':['checkout','process','coupon','mobile','finding','guide','confusing'],
    'quality':['perfect','good','quality','pictured','product'],
    'experience':['pleasant','recommend','shopping','experience']
}

def extract_theme_adv(text):
    found = []
    for t,words in themes_advanced.items():
        if any(w in text for w in words):
            found.append(t)
    return ', '.join(found) if found else 'other'

real_df['Theme'] = real_df['Feedback_clean'].apply(extract_theme_adv)
print(real_df[['Feedback','Theme']])
                                            Feedback                Theme
0  The checkout process was confusing and I could...              website
1  Shipping was extremely fast and everything arr...    shipping, quality
2   I had trouble finding the sizing guide on mobile              website
3           Would recommend. Good quality and price!  quality, experience
4     My issue took too long to resolve with support              support
5                     The product is not as pictured              quality
6                  Very pleasant shopping experience           experience
7               Got my refund right away, thank you!              support
theme_dist = real_df['Theme'].str.get_dummies(sep=', ').sum(axis=0).sort_values(ascending=False)
print(theme_dist)
quality       3
experience    2
support       2
website       2
shipping      1
dtype: int64
theme_dist.plot(kind='bar',title='Theme Mentions in Customer Comments',color='steelblue')
plt.ylabel('Mentions')
plt.show()
No description has been provided for this image

Error handling and debugging (missing feedback)#

  • Real-world survey data often has missing written comments.
  • We must detect and handle missing or empty feedback to avoid bias or errors.
test_df = pd.DataFrame({
    'CustomerID':[11,12,13],
    'Feedback':['Great service',np.nan,'']
})
missing = test_df['Feedback'].isnull() | (test_df['Feedback'].str.strip() == '')
print('Missing feedback entries:')
print(test_df[missing])
Missing feedback entries:
   CustomerID Feedback
1          12      NaN
2          13         
sample_df = pd.DataFrame({'Theme':['service','delivery','service','value','service']})
wrong_counts = sample_df.groupby('Theme').count()
print('Incorrect grouping result:')
print(wrong_counts)

right_counts = sample_df.groupby('Theme').size()
print('Correct theme counts:')
print(right_counts)
Incorrect grouping result:
Empty DataFrame
Columns: []
Index: [delivery, service, value]
Correct theme counts:
Theme
delivery    1
service     3
value       1
dtype: int64
likert_scores = pd.Series(['Strongly agree', 'Agree', 'Neutral', 'Disagree', 'Strongly agree'])
score_map = {
    'Strongly agree':5,
    'Agree':4,
    'Neutral':3,
    'Disagree':2,
    'Strongly disagree':1
}
num_scores = likert_scores.map(score_map)
print(num_scores)
0    5
1    4
2    3
3    2
4    5
dtype: int64

Best practices for analyzing open feedback#

  • Segment comments by customer type or sentiment, not just topic.
  • Cross-tabulate themes with NPS groups, age, or purchase size.
  • Create composite indices or theme scores if management wants summary KPIs.
  • Track trends in complaint or praise themes over time for early warnings.
# Example: Cross-tabulate theme by NPS group from earlier
print(result.groupby(['Theme','NPS_Group']).size().unstack().fillna(0))
NPS_Group             Detractor  Passive  Promoter
Theme                                             
delivery                    1.0      0.0       0.0
quality                     0.0      0.0       1.0
quality, value              0.0      1.0       0.0
service                     0.0      0.0       1.0
service, improvement        1.0      0.0       0.0
# Build a theme score (1 point for each theme per comment)
result['Theme_Score'] = result['Theme'].apply(lambda x: 0 if x=='other' else len(x.split(',')))
print(result[['Feedback','Theme','Theme_Score']])
                                   Feedback                 Theme  Theme_Score
0          Great service and friendly staff               service            1
1  Delivery was slow and packaging was poor              delivery            1
2         Excellent quality, will buy again               quality            1
3        Customer support needs improvement  service, improvement            2
4                      Good value for money        quality, value            2
# Example: Trend analysis (simulate monthly feedback counts by theme)
trend_df = pd.DataFrame({
    'Month':['2023-01','2023-01','2023-02','2023-02','2023-02','2023-03','2023-03'],
    'Theme':['service','delivery','service','delivery','value','value','service']
})
pivot = trend_df.pivot_table(index='Month',columns='Theme',values='Theme',aggfunc='count').fillna(0)
print(pivot)
Empty DataFrame
Columns: []
Index: [2023-01, 2023-02, 2023-03]

End-to-end market research insight: What can we recommend?#

  • Collect raw feedback data and assign business themes using keywords.
  • Link themes to customer advocacy or satisfaction groups.
  • Compute the most common complaint and praise topics.
  • Recommend actionable changes where themes show high volume or negative sentiment.
  • Summarize findings for business teams in charts and tables.
# Workflow: From raw data to recommended business actions
final_theme_counts = real_df['Theme'].str.get_dummies(sep=', ').sum(axis=0)
main_problem = final_theme_counts.idxmax()
print('The most common feedback theme is:', main_problem)
if main_problem in ['shipping','support','website']:
    print('Business should review processes in', main_problem, 'immediately!')
The most common feedback theme is: quality
 

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.