Lesson 45 · Market Research Analytics in Python
Visualizing Text-Based Market Insights in Python for Market Research
Market researchers often receive open-ended customer feedback as unstructured text. Understanding customer sentiments, concerns, and suggestions from this…
- CourseMarket Research Analytics in Python
- Lesson45 of 56
- Video22 min
- FormatJupyter notebook · 21 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbVisualizing Text-Based Market Insights#
- Market researchers often receive open-ended customer feedback as unstructured text.
- Understanding customer sentiments, concerns, and suggestions from this text can help businesses improve products and services.
- In this lesson, you will learn to visualize text-based market insights using real survey and feedback data.
- We will explore how to extract meaning, spot trends, and present findings in clear visual forms.
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns
from wordcloud import WordCloud, STOPWORDS
import warnings
warnings.filterwarnings('ignore')
Understanding Survey and Feedback Data#
- Surveys and customer feedback data often contain both structured (ratings, choices) and unstructured (written comments) responses.
- Text columns capture nuanced opinions that are not visible in numeric scores alone.
- Typical columns: customer ID, feedback text, survey score, demographic info.
- Beginners often forget to clean text, handle missing values, or ignore the importance of open text fields.
- Always review a sample of raw comments to gain context before analysis.
# Load a sample open-ended customer feedback dataset
df = pd.DataFrame({'CustomerID':[1,2,3,4,5],
'Feedback':['Great service and friendly staff',
'Delivery was slow and packaging was poor',
'Excellent quality, will buy again',
'Customer support needs improvement',
'Good value for money']})
print(df.shape)
print(df.head())
# See feedback comments quickly
for idx, row in df.iterrows():
print(f'Customer {row.CustomerID}: "{row.Feedback}"')
# Word frequency counts: simple example
from collections import Counter
feedback_text = ' '.join(df['Feedback']).lower()
words = feedback_text.split()
stopwords = set(STOPWORDS)
filtered_words = [w for w in words if w not in stopwords]
word_counts = Counter(filtered_words)
print(word_counts.most_common(7))
# Basic word cloud visualization
wc = WordCloud(width=480, height=240, background_color='white', stopwords=STOPWORDS).generate(feedback_text)
plt.figure(figsize=(8,4))
plt.imshow(wc, interpolation='bilinear')
plt.axis('off')
plt.title('Most Common Words in Customer Feedback')
plt.show()
# Add a longer feedback set for richer analysis
more_data = [
'Staff did not help me with my question',
'Fast shipping and clean packaging',
'I found the website confusing',
'Always a great deal at this store',
'Delivery timing could be better',
'Returns process was simple',
'The product broke after one week',
'Excellent value and quick response from support'
]
more_ids = range(6, 14)
df2 = pd.DataFrame({'CustomerID': list(more_ids), 'Feedback': more_data})
df = pd.concat([df, df2], ignore_index=True)
print(df.shape)
print(df.tail())
# Create a word cloud from the larger feedback set
feedback_text = ' '.join(df['Feedback']).lower()
wc = WordCloud(width=520, height=260, background_color='white', stopwords=STOPWORDS, max_words=30).generate(feedback_text)
plt.figure(figsize=(10,5))
plt.imshow(wc, interpolation='bilinear')
plt.axis('off')
plt.title('Customer Feedback: Top 30 Words')
plt.show()
# Bar plot of most common feedback words
from collections import Counter
words = feedback_text.split()
filtered_words = [w for w in words if w not in stopwords]
top_words = Counter(filtered_words).most_common(10)
words, counts = zip(*top_words)
plt.figure(figsize=(8,4))
sns.barplot(x=list(counts), y=list(words), palette='Blues_r')
plt.xlabel('Count')
plt.title('Top 10 Words in Customer Feedback')
plt.tight_layout()
plt.show()
# Adding sentiment: quick manual tagging
def tag_sentiment(text):
positives = ['great', 'excellent', 'good', 'fast', 'clean', 'value', 'simple', 'quick', 'help']
negatives = ['poor', 'slow', 'broke', 'confusing', 'did not', 'needs', 'not']
text_lower = text.lower()
if any(word in text_lower for word in positives):
return 'Positive'
elif any(word in text_lower for word in negatives):
return 'Negative'
else:
return 'Neutral'
df['Sentiment'] = df['Feedback'].apply(tag_sentiment)
print(df[['Feedback', 'Sentiment']].head(8))
# Visualize distribution of sentiments
sent_counts = df['Sentiment'].value_counts()
sns.barplot(x=sent_counts.index, y=sent_counts.values, palette='viridis')
plt.xlabel('Sentiment')
plt.ylabel('Number of Comments')
plt.title('Feedback Sentiment Distribution')
plt.tight_layout()
plt.show()
# Feedback themes by sentiment: intermediate pattern
for sentiment in ['Positive','Negative','Neutral']:
subset = df[df['Sentiment']==sentiment]
if len(subset):
subset_text = ' '.join(subset['Feedback']).lower()
sub_wc = WordCloud(width=420, height=190, background_color='white', stopwords=STOPWORDS, max_words=10).generate(subset_text)
plt.figure(figsize=(6,3))
plt.imshow(sub_wc, interpolation='bilinear')
plt.axis('off')
plt.title(f'{sentiment} Feedback: Top Words')
plt.show()
# Connecting text feedback to scores: setup NPS example
np.random.seed(42)
nps_scores = np.random.randint(0, 11, size=len(df))
df['NPS_Score'] = nps_scores
print(df[['Feedback','NPS_Score']].head())
# Visualize average NPS by sentiment
nps_by_sent = df.groupby('Sentiment')['NPS_Score'].mean()
sns.barplot(x=nps_by_sent.index, y=nps_by_sent.values, palette='mako')
plt.ylabel('Average NPS Score')
plt.title('Average NPS by Feedback Sentiment')
plt.tight_layout()
plt.show()
# Intermediate: Find and visualize top trigrams in feedback
from sklearn.feature_extraction.text import CountVectorizer
vectorizer = CountVectorizer(stop_words='english', ngram_range=(3,3))
X = vectorizer.fit_transform(df['Feedback'])
sum_words = X.sum(axis=0)
words_freq = [(word, sum_words[0, idx]) for word, idx in vectorizer.vocabulary_.items()]
words_freq = sorted(words_freq, key=lambda x: x[1], reverse=True)[:7]
trigrams, counts = zip(*words_freq) if words_freq else ([], [])
plt.figure(figsize=(10,3))
sns.barplot(x=list(counts), y=list(trigrams), palette='flare')
plt.xlabel('Count')
plt.title('Most Common Trigrams in Feedback')
plt.tight_layout()
plt.show()
# Advanced: Connect text feedback to customer demographics
ages = np.random.randint(18, 65, len(df))
regions = np.random.choice(['North', 'South', 'East', 'West'], len(df))
df['Age'] = ages
df['Region'] = regions
plt.figure(figsize=(7,4))
sns.boxplot(x='Region', y='NPS_Score', data=df, palette='pastel')
plt.title('NPS Score Distribution by Region')
plt.show()
# Advanced: Topic separation using simple keyword matching
df['Topic'] = df['Feedback'].str.contains('delivery|shipping|packaging|timing', case=False).map({True: 'Delivery', False: 'Other'})
topic_counts = df.groupby(['Topic','Sentiment']).size().unstack(fill_value=0)
topic_counts.plot(kind='bar', stacked=True, color=['green','red','grey'], figsize=(7,4))
plt.title('Topic vs Sentiment in Feedback')
plt.ylabel('Number of Comments')
plt.xticks(rotation=0)
plt.tight_layout()
plt.show()
# Error handling 1: Missing comment fills
df_with_missing = df.copy()
df_with_missing.loc[2,'Feedback'] = np.nan
df_with_missing['Feedback_filled'] = df_with_missing['Feedback'].fillna('no feedback provided')
print(df_with_missing[['CustomerID','Feedback','Feedback_filled']].head(5))
# Error handling 2: Incorrect groupings
try:
# Wrong: Group by non-existing column
test = df.groupby('Sent').size()
except Exception as e:
print('Grouping error:', e)
# Error handling 3: Misinterpreting sentiment and NPS
df['Likert_Score'] = np.where(df['NPS_Score'] >= 8, 'Agree', np.where(df['NPS_Score'] <= 3, 'Disagree', 'Neutral'))
crosstab = pd.crosstab(df['Sentiment'], df['Likert_Score'])
print(crosstab)
Best Practices: Patterns for Text Market Research#
- Segment feedback by meaningful groups: product, location, score tier.
- Always remove stopwords and check for spelling variations.
- Use cross-tabs to relate text sentiment to numeric ratings.
- Build indexes or thematic scores by counting major topics.
- Visualize trends over time if you have dates.
# Pattern: Index construction (complaints index)
df['Is_Complaint'] = df['Feedback'].str.contains('broke|did not|needs|poor|confusing|not help', case=False)
complaint_index = df['Is_Complaint'].mean()
print(f'Complaint Index: {complaint_index:.2%} of comments are complaints')
# Pattern: Follow-up deep dive on frequent complaint
complaints = df[df['Is_Complaint']]
print('Sample complaints:')
print(complaints['Feedback'].tolist()[:5])
End-to-End Example: Insight to Recommendation#
- We start with customer feedback text and basic demographic segmentation.
- We identify dominant positive and negative topics using word clouds and bar plots.
- Sentiment analysis and complaint indexing provide numeric metrics.
- We connect these insights to NPS for actionable reporting.
- Recommendations: Focus improvements on most-frequent complaint topics and regions with lowest average NPS.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



