Mathew K Analytics

Lesson 42 · Market Research Analytics in Python

Master Word Frequency & Keyword Analysis for Market Research Using Python

In this lesson, we will explore how to analyze customer feedback for frequent words and keywords. This helps uncover patterns, highlight major issues, and…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Word Frequency and Keyword Analysis in Customer Feedback#

  • In this lesson, we will explore how to analyze customer feedback for frequent words and keywords.
  • This helps uncover patterns, highlight major issues, and identify opportunities in customer experience.
  • Learning to extract word frequencies helps businesses prioritize improvements based on real customer voice.
  • You will perform these analyses on real and realistic datasets, interpreting outputs to drive actionable insights.
import pandas as pd
import numpy as np
import re
from collections import Counter
import warnings
warnings.filterwarnings('ignore')

Understanding the Data: Market Research and Customer Analytics#

  • We work with open-ended feedback and survey data from real and synthetic datasets.
  • Each row usually represents one customer's response or piece of feedback.
  • Feedback may include text (open-ended), ratings, demographics, or both.
  • Common beginner mistakes include ignoring text cleaning, forgetting to lower case, double-counting, and misinterpreting word importance.
df = pd.DataFrame({'CustomerID':[1,2,3,4,5],
                  'Feedback':['Great service and friendly staff','Delivery was slow and packaging was poor','Excellent quality, will buy again','Customer support needs improvement','Good value for money']})
print(df.shape)
print(df.head(3))
(5, 2)
   CustomerID                                  Feedback
0           1          Great service and friendly staff
1           2  Delivery was slow and packaging was poor
2           3         Excellent quality, will buy again
all_feedback = ' '.join(df['Feedback'])
print(all_feedback)
Great service and friendly staff Delivery was slow and packaging was poor Excellent quality, will buy again Customer support needs improvement Good value for money
words = re.findall(r'\b\w+\b', all_feedback.lower())
print(words)
['great', 'service', 'and', 'friendly', 'staff', 'delivery', 'was', 'slow', 'and', 'packaging', 'was', 'poor', 'excellent', 'quality', 'will', 'buy', 'again', 'customer', 'support', 'needs', 'improvement', 'good', 'value', 'for', 'money']
word_counts = Counter(words)
print(word_counts)
Counter({'and': 2, 'was': 2, 'great': 1, 'service': 1, 'friendly': 1, 'staff': 1, 'delivery': 1, 'slow': 1, 'packaging': 1, 'poor': 1, 'excellent': 1, 'quality': 1, 'will': 1, 'buy': 1, 'again': 1, 'customer': 1, 'support': 1, 'needs': 1, 'improvement': 1, 'good': 1, 'value': 1, 'for': 1, 'money': 1})
print('Most common words:')
print(word_counts.most_common(5))
Most common words:
[('and', 2), ('was', 2), ('great', 1), ('service', 1), ('friendly', 1)]
stopwords = {'and', 'was', 'for', 'the', 'will', 'again'}
words_filtered = [w for w in words if w not in stopwords]
filtered_counts = Counter(words_filtered)
print(filtered_counts.most_common(5))
[('great', 1), ('service', 1), ('friendly', 1), ('staff', 1), ('delivery', 1)]
# Beginner: Single feedback word count
feedback_row = df.iloc[2]['Feedback']
words_row = re.findall(r'\b\w+\b', feedback_row.lower())
row_counts = Counter(words_row)
print(f'Most frequent word(s) in one feedback: {row_counts.most_common(2)}')
Most frequent word(s) in one feedback: [('excellent', 1), ('quality', 1)]
# Beginner: Count positive sentiment keywords
positive_words = {'great', 'excellent', 'good', 'friendly', 'quality', 'value'}
positive_in_feedback = [w for w in words if w in positive_words]
print(f'Positive keywords appeared {len(positive_in_feedback)} times.')
Positive keywords appeared 6 times.
# Beginner: Find feedbacks containing 'slow' or 'support'
keyword = 'slow'
rows_with_keyword = df[df['Feedback'].str.lower().str.contains(keyword)]
print(rows_with_keyword)
   CustomerID                                  Feedback
1           2  Delivery was slow and packaging was poor
# Intermediate: Apply word frequency to a larger dataset
import openml
dataset = openml.datasets.get_dataset(42178)
df_survey, _, _, _ = dataset.get_data(dataset_format='dataframe')
feedback_col = 'Churn'
feedback_texts = df['Feedback'].tolist() + ['satisfied' if val == 'No' else 'dissatisfied' for val in df_survey[feedback_col][:5]]
all_feedback_large = ' '.join([str(fb) for fb in feedback_texts]).lower()
words_large = re.findall(r'\b\w+\b', all_feedback_large)
counts_large = Counter(words_large)
print(counts_large.most_common(8))
[('satisfied', 3), ('and', 2), ('was', 2), ('dissatisfied', 2), ('great', 1), ('service', 1), ('friendly', 1), ('staff', 1)]
# Intermediate: Keyword analysis by customer type
sample_df = df.copy()
sample_df['Type'] = ['New', 'Returning', 'Returning', 'New', 'New']
for cat in sample_df['Type'].unique():
    cat_feedback = ' '.join(sample_df[sample_df['Type'] == cat]['Feedback']).lower()
    cat_words = re.findall(r'\b\w+\b', cat_feedback)
    print(f'Customer type: {cat}')
    print(Counter(cat_words).most_common(3))
Customer type: New
[('great', 1), ('service', 1), ('and', 1)]
Customer type: Returning
[('was', 2), ('delivery', 1), ('slow', 1)]
# Intermediate: Cross-tab of keyword vs. customer segment
results = []
keywords = ['slow', 'value', 'support', 'quality']
for word in keywords:
    for t in sample_df['Type'].unique():
        txt = ' '.join(sample_df[sample_df['Type']==t]['Feedback']).lower()
        count = txt.count(word)
        results.append({'Keyword':word, 'Type':t, 'Count':count})
ct = pd.DataFrame(results).pivot(index='Keyword', columns='Type', values='Count')
print(ct)
Type     New  Returning
Keyword                
quality    0          1
slow       0          1
support    1          0
value      1          0
# Advanced: Extract bigrams commonly used by customers
def bigrams(wordlist):
    return [' '.join([wordlist[i], wordlist[i+1]]) for i in range(len(wordlist)-1)]
bigrams_all = bigrams(words_filtered)
bigram_counts = Counter(bigrams_all)
print('Most common bigrams:', bigram_counts.most_common(3))
Most common bigrams: [('great service', 1), ('service friendly', 1), ('friendly staff', 1)]
# Advanced: Word cloud visualization for management report
from wordcloud import WordCloud
import matplotlib.pyplot as plt
wc = WordCloud(width=400, height=200, background_color='white', collocations=False)
wc.generate(' '.join(words_filtered))
plt.figure(figsize=(6,3))
plt.imshow(wc, interpolation='bilinear')
plt.axis('off')
plt.title('Customer Feedback Word Cloud')
plt.tight_layout()
plt.savefig('customer_feedback_wordcloud.png')
plt.show()
No description has been provided for this image
# Advanced: Keyword trends over time (simulated timestamps)
import random
random.seed(42)
sample_df['Date'] = pd.date_range('2023-01-01', periods=len(sample_df))
sample_df.loc[2, 'Feedback'] = 'Quality is good but support is slow'
sample_df.loc[3, 'Feedback'] = 'Service was great, packaging was poor'
trend_kw = 'support'
sample_df['Mentions'] = sample_df['Feedback'].str.lower().apply(lambda x: trend_kw in x)
grouped = sample_df.groupby(sample_df['Date'].dt.month)['Mentions'].sum()
print('Monthly mentions of the word "support":')
print(grouped)
Monthly mentions of the word "support":
Date
1    1
Name: Mentions, dtype: int64
# Error handling: Handling missing feedback responses
df_missing = df.copy()
df_missing.loc[1, 'Feedback'] = None
print('Before fill:', df_missing['Feedback'].tolist())
df_missing['Feedback'] = df_missing['Feedback'].fillna('No feedback provided')
print('After fill:', df_missing['Feedback'].tolist())
Before fill: ['Great service and friendly staff', None, 'Excellent quality, will buy again', 'Customer support needs improvement', 'Good value for money']
After fill: ['Great service and friendly staff', 'No feedback provided', 'Excellent quality, will buy again', 'Customer support needs improvement', 'Good value for money']
# Error handling: Incorrect groupings in keyword counts
try:
    # Misgroup: grouping on the wrong column
    wrg = df.groupby('CustomerID')['Feedback'].apply(lambda x: 'slow' in x.astype(str).str.lower().str.cat(sep=' ')).sum()
    print('Wrong sum:', wrg)
except Exception as e:
    print('Error caught:', e)
# Correct approach:
num_mention = df['Feedback'].str.lower().str.contains('slow').sum()
print('Number of feedbacks mentioning "slow":', num_mention)
Wrong sum: 1
Number of feedbacks mentioning "slow": 1
# Error handling: Misinterpreting NPS scores as text
import numpy as np
np.random.seed(42)
nps_df = pd.DataFrame({'CustomerID': range(1,11),
                      'Feedback': ['Promoter', 'Promoter', 'Detractor', 'Passive', 'Passive', 'Detractor', 'Promoter', 'Detractor', 'Promoter', 'Promoter'],
                      'NPS_Score': np.random.randint(0,11,10)})
nps_df['Feedback'] = nps_df['Feedback'].astype(str)
promoters = nps_df[nps_df['NPS_Score'] >= 9]
print(f'Number of promoters (by score): {len(promoters)}')
Number of promoters (by score): 3

Best Practices and Patterns: Market Research Keyword Analysis#

  • Always review and clean text for irrelevant words (stopwords, typos, filler).
  • Segment feedback by customer type, time, channel, or product for deeper insights.
  • Cross-tabulate keywords with demographic or behavioral attributes.
  • Build indexes for satisfaction or complaint volume by keyword.
  • Monitor trends of keywords across months, campaigns, or market events.
# Pattern: Segmentation by region in NPS survey data
np.random.seed(42)
nps_large = pd.DataFrame({
    'CustomerID': range(1,101),
    'Region': np.random.choice(['North','South','East','West'], 100),
    'NPS_Score': np.random.randint(0,11,100),
    'Feedback': np.random.choice(['Great value', 'Very slow service', 'Excellent support', 'Average', 'Happy customer'], 100)
})
for reg in nps_large['Region'].unique():
    words_reg = re.findall(r'\b\w+\b', ' '.join(nps_large[nps_large['Region']==reg]['Feedback']).lower())
    print(f'Most common in {reg}:', Counter(words_reg).most_common(2))
Most common in East: [('excellent', 10), ('support', 10)]
Most common in West: [('great', 7), ('value', 7)]
Most common in North: [('great', 5), ('value', 5)]
Most common in South: [('average', 10), ('great', 8)]
# Pattern: Construct an index of negative vs. positive keyword balance
keywords_pos = ['great', 'excellent', 'happy', 'support', 'value']
keywords_neg = ['slow', 'poor', 'average']
idx = 0
for fb in df['Feedback'].str.lower():
    wc = Counter(re.findall(r'\b\w+\b', fb))
    idx += sum([wc[w] for w in keywords_pos if w in wc])
    idx -= sum([wc[w] for w in keywords_neg if w in wc])
print('Net word sentiment index:', idx)
Net word sentiment index: 2
# Tiny end-to-end word frequency problem: Find most common complaint in campaign data
import openml
dataset = openml.datasets.get_dataset(1461)
df_campaign, _, _, _ = dataset.get_data(dataset_format='dataframe')
df_campaign.columns = ['age','job','marital','education','default','balance','housing','loan','contact','day','month','duration','campaign','pdays','previous','poutcome','response']
complaints = [
    'I was not informed about changes', 'Slow response from staff', 'Good service overall',
    'The offer was unclear', 'Delayed processing', 'Fast and friendly', 'Did not resolve my issue',
    'Good experience', 'Late follow-up', 'Excellent support'
]
np.random.seed(42)
df_campaign['Feedback'] = np.random.choice(complaints, len(df_campaign))
wordlist = re.findall(r'\b\w+\b', ' '.join(df_campaign['Feedback']).lower())
complaint_words = ['slow', 'late', 'not', 'did', 'unclear', 'delayed', 'issue']
counts_complaint = {w: wordlist.count(w) for w in complaint_words}
print('Complaint keyword frequencies:', counts_complaint)
main_issue = max(counts_complaint, key=counts_complaint.get)
print(f'Top recurring complaint term in marketing campaign: {main_issue}')
Complaint keyword frequencies: {'slow': 4443, 'late': 4484, 'not': 9063, 'did': 4610, 'unclear': 4548, 'delayed': 4490, 'issue': 4610}
Top recurring complaint term in marketing campaign: not
 

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.