Lesson 42 · Market Research Analytics in Python
Master Word Frequency & Keyword Analysis for Market Research Using Python
In this lesson, we will explore how to analyze customer feedback for frequent words and keywords. This helps uncover patterns, highlight major issues, and…
- CourseMarket Research Analytics in Python
- Lesson42 of 56
- Video24 min
- FormatJupyter notebook · 23 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbWord Frequency and Keyword Analysis in Customer Feedback#
- In this lesson, we will explore how to analyze customer feedback for frequent words and keywords.
- This helps uncover patterns, highlight major issues, and identify opportunities in customer experience.
- Learning to extract word frequencies helps businesses prioritize improvements based on real customer voice.
- You will perform these analyses on real and realistic datasets, interpreting outputs to drive actionable insights.
import pandas as pd
import numpy as np
import re
from collections import Counter
import warnings
warnings.filterwarnings('ignore')
Understanding the Data: Market Research and Customer Analytics#
- We work with open-ended feedback and survey data from real and synthetic datasets.
- Each row usually represents one customer's response or piece of feedback.
- Feedback may include text (open-ended), ratings, demographics, or both.
- Common beginner mistakes include ignoring text cleaning, forgetting to lower case, double-counting, and misinterpreting word importance.
df = pd.DataFrame({'CustomerID':[1,2,3,4,5],
'Feedback':['Great service and friendly staff','Delivery was slow and packaging was poor','Excellent quality, will buy again','Customer support needs improvement','Good value for money']})
print(df.shape)
print(df.head(3))
all_feedback = ' '.join(df['Feedback'])
print(all_feedback)
words = re.findall(r'\b\w+\b', all_feedback.lower())
print(words)
word_counts = Counter(words)
print(word_counts)
print('Most common words:')
print(word_counts.most_common(5))
stopwords = {'and', 'was', 'for', 'the', 'will', 'again'}
words_filtered = [w for w in words if w not in stopwords]
filtered_counts = Counter(words_filtered)
print(filtered_counts.most_common(5))
# Beginner: Single feedback word count
feedback_row = df.iloc[2]['Feedback']
words_row = re.findall(r'\b\w+\b', feedback_row.lower())
row_counts = Counter(words_row)
print(f'Most frequent word(s) in one feedback: {row_counts.most_common(2)}')
# Beginner: Count positive sentiment keywords
positive_words = {'great', 'excellent', 'good', 'friendly', 'quality', 'value'}
positive_in_feedback = [w for w in words if w in positive_words]
print(f'Positive keywords appeared {len(positive_in_feedback)} times.')
# Beginner: Find feedbacks containing 'slow' or 'support'
keyword = 'slow'
rows_with_keyword = df[df['Feedback'].str.lower().str.contains(keyword)]
print(rows_with_keyword)
# Intermediate: Apply word frequency to a larger dataset
import openml
dataset = openml.datasets.get_dataset(42178)
df_survey, _, _, _ = dataset.get_data(dataset_format='dataframe')
feedback_col = 'Churn'
feedback_texts = df['Feedback'].tolist() + ['satisfied' if val == 'No' else 'dissatisfied' for val in df_survey[feedback_col][:5]]
all_feedback_large = ' '.join([str(fb) for fb in feedback_texts]).lower()
words_large = re.findall(r'\b\w+\b', all_feedback_large)
counts_large = Counter(words_large)
print(counts_large.most_common(8))
# Intermediate: Keyword analysis by customer type
sample_df = df.copy()
sample_df['Type'] = ['New', 'Returning', 'Returning', 'New', 'New']
for cat in sample_df['Type'].unique():
cat_feedback = ' '.join(sample_df[sample_df['Type'] == cat]['Feedback']).lower()
cat_words = re.findall(r'\b\w+\b', cat_feedback)
print(f'Customer type: {cat}')
print(Counter(cat_words).most_common(3))
# Intermediate: Cross-tab of keyword vs. customer segment
results = []
keywords = ['slow', 'value', 'support', 'quality']
for word in keywords:
for t in sample_df['Type'].unique():
txt = ' '.join(sample_df[sample_df['Type']==t]['Feedback']).lower()
count = txt.count(word)
results.append({'Keyword':word, 'Type':t, 'Count':count})
ct = pd.DataFrame(results).pivot(index='Keyword', columns='Type', values='Count')
print(ct)
# Advanced: Extract bigrams commonly used by customers
def bigrams(wordlist):
return [' '.join([wordlist[i], wordlist[i+1]]) for i in range(len(wordlist)-1)]
bigrams_all = bigrams(words_filtered)
bigram_counts = Counter(bigrams_all)
print('Most common bigrams:', bigram_counts.most_common(3))
# Advanced: Word cloud visualization for management report
from wordcloud import WordCloud
import matplotlib.pyplot as plt
wc = WordCloud(width=400, height=200, background_color='white', collocations=False)
wc.generate(' '.join(words_filtered))
plt.figure(figsize=(6,3))
plt.imshow(wc, interpolation='bilinear')
plt.axis('off')
plt.title('Customer Feedback Word Cloud')
plt.tight_layout()
plt.savefig('customer_feedback_wordcloud.png')
plt.show()
# Advanced: Keyword trends over time (simulated timestamps)
import random
random.seed(42)
sample_df['Date'] = pd.date_range('2023-01-01', periods=len(sample_df))
sample_df.loc[2, 'Feedback'] = 'Quality is good but support is slow'
sample_df.loc[3, 'Feedback'] = 'Service was great, packaging was poor'
trend_kw = 'support'
sample_df['Mentions'] = sample_df['Feedback'].str.lower().apply(lambda x: trend_kw in x)
grouped = sample_df.groupby(sample_df['Date'].dt.month)['Mentions'].sum()
print('Monthly mentions of the word "support":')
print(grouped)
# Error handling: Handling missing feedback responses
df_missing = df.copy()
df_missing.loc[1, 'Feedback'] = None
print('Before fill:', df_missing['Feedback'].tolist())
df_missing['Feedback'] = df_missing['Feedback'].fillna('No feedback provided')
print('After fill:', df_missing['Feedback'].tolist())
# Error handling: Incorrect groupings in keyword counts
try:
# Misgroup: grouping on the wrong column
wrg = df.groupby('CustomerID')['Feedback'].apply(lambda x: 'slow' in x.astype(str).str.lower().str.cat(sep=' ')).sum()
print('Wrong sum:', wrg)
except Exception as e:
print('Error caught:', e)
# Correct approach:
num_mention = df['Feedback'].str.lower().str.contains('slow').sum()
print('Number of feedbacks mentioning "slow":', num_mention)
# Error handling: Misinterpreting NPS scores as text
import numpy as np
np.random.seed(42)
nps_df = pd.DataFrame({'CustomerID': range(1,11),
'Feedback': ['Promoter', 'Promoter', 'Detractor', 'Passive', 'Passive', 'Detractor', 'Promoter', 'Detractor', 'Promoter', 'Promoter'],
'NPS_Score': np.random.randint(0,11,10)})
nps_df['Feedback'] = nps_df['Feedback'].astype(str)
promoters = nps_df[nps_df['NPS_Score'] >= 9]
print(f'Number of promoters (by score): {len(promoters)}')
Best Practices and Patterns: Market Research Keyword Analysis#
- Always review and clean text for irrelevant words (stopwords, typos, filler).
- Segment feedback by customer type, time, channel, or product for deeper insights.
- Cross-tabulate keywords with demographic or behavioral attributes.
- Build indexes for satisfaction or complaint volume by keyword.
- Monitor trends of keywords across months, campaigns, or market events.
# Pattern: Segmentation by region in NPS survey data
np.random.seed(42)
nps_large = pd.DataFrame({
'CustomerID': range(1,101),
'Region': np.random.choice(['North','South','East','West'], 100),
'NPS_Score': np.random.randint(0,11,100),
'Feedback': np.random.choice(['Great value', 'Very slow service', 'Excellent support', 'Average', 'Happy customer'], 100)
})
for reg in nps_large['Region'].unique():
words_reg = re.findall(r'\b\w+\b', ' '.join(nps_large[nps_large['Region']==reg]['Feedback']).lower())
print(f'Most common in {reg}:', Counter(words_reg).most_common(2))
# Pattern: Construct an index of negative vs. positive keyword balance
keywords_pos = ['great', 'excellent', 'happy', 'support', 'value']
keywords_neg = ['slow', 'poor', 'average']
idx = 0
for fb in df['Feedback'].str.lower():
wc = Counter(re.findall(r'\b\w+\b', fb))
idx += sum([wc[w] for w in keywords_pos if w in wc])
idx -= sum([wc[w] for w in keywords_neg if w in wc])
print('Net word sentiment index:', idx)
# Tiny end-to-end word frequency problem: Find most common complaint in campaign data
import openml
dataset = openml.datasets.get_dataset(1461)
df_campaign, _, _, _ = dataset.get_data(dataset_format='dataframe')
df_campaign.columns = ['age','job','marital','education','default','balance','housing','loan','contact','day','month','duration','campaign','pdays','previous','poutcome','response']
complaints = [
'I was not informed about changes', 'Slow response from staff', 'Good service overall',
'The offer was unclear', 'Delayed processing', 'Fast and friendly', 'Did not resolve my issue',
'Good experience', 'Late follow-up', 'Excellent support'
]
np.random.seed(42)
df_campaign['Feedback'] = np.random.choice(complaints, len(df_campaign))
wordlist = re.findall(r'\b\w+\b', ' '.join(df_campaign['Feedback']).lower())
complaint_words = ['slow', 'late', 'not', 'did', 'unclear', 'delayed', 'issue']
counts_complaint = {w: wordlist.count(w) for w in complaint_words}
print('Complaint keyword frequencies:', counts_complaint)
main_issue = max(counts_complaint, key=counts_complaint.get)
print(f'Top recurring complaint term in marketing campaign: {main_issue}')
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



