Mathew K Analytics

Lesson 45 · Market Research Analytics in Python

Visualizing Text-Based Market Insights in Python for Market Research

Market researchers often receive open-ended customer feedback as unstructured text. Understanding customer sentiments, concerns, and suggestions from this…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Visualizing Text-Based Market Insights#

  • Market researchers often receive open-ended customer feedback as unstructured text.
  • Understanding customer sentiments, concerns, and suggestions from this text can help businesses improve products and services.
  • In this lesson, you will learn to visualize text-based market insights using real survey and feedback data.
  • We will explore how to extract meaning, spot trends, and present findings in clear visual forms.
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns
from wordcloud import WordCloud, STOPWORDS
import warnings
warnings.filterwarnings('ignore')

Understanding Survey and Feedback Data#

  • Surveys and customer feedback data often contain both structured (ratings, choices) and unstructured (written comments) responses.
  • Text columns capture nuanced opinions that are not visible in numeric scores alone.
  • Typical columns: customer ID, feedback text, survey score, demographic info.
  • Beginners often forget to clean text, handle missing values, or ignore the importance of open text fields.
  • Always review a sample of raw comments to gain context before analysis.
# Load a sample open-ended customer feedback dataset
df = pd.DataFrame({'CustomerID':[1,2,3,4,5],
    'Feedback':['Great service and friendly staff',
               'Delivery was slow and packaging was poor',
               'Excellent quality, will buy again',
               'Customer support needs improvement',
               'Good value for money']})
print(df.shape)
print(df.head())
(5, 2)
   CustomerID                                  Feedback
0           1          Great service and friendly staff
1           2  Delivery was slow and packaging was poor
2           3         Excellent quality, will buy again
3           4        Customer support needs improvement
4           5                      Good value for money
# See feedback comments quickly
for idx, row in df.iterrows():
    print(f'Customer {row.CustomerID}: "{row.Feedback}"')
Customer 1: "Great service and friendly staff"
Customer 2: "Delivery was slow and packaging was poor"
Customer 3: "Excellent quality, will buy again"
Customer 4: "Customer support needs improvement"
Customer 5: "Good value for money"
# Word frequency counts: simple example
from collections import Counter
feedback_text = ' '.join(df['Feedback']).lower()
words = feedback_text.split()
stopwords = set(STOPWORDS)
filtered_words = [w for w in words if w not in stopwords]
word_counts = Counter(filtered_words)
print(word_counts.most_common(7))
[('great', 1), ('service', 1), ('friendly', 1), ('staff', 1), ('delivery', 1), ('slow', 1), ('packaging', 1)]
# Basic word cloud visualization
wc = WordCloud(width=480, height=240, background_color='white', stopwords=STOPWORDS).generate(feedback_text)
plt.figure(figsize=(8,4))
plt.imshow(wc, interpolation='bilinear')
plt.axis('off')
plt.title('Most Common Words in Customer Feedback')
plt.show()
No description has been provided for this image
# Add a longer feedback set for richer analysis
more_data = [
    'Staff did not help me with my question',
    'Fast shipping and clean packaging',
    'I found the website confusing',
    'Always a great deal at this store',
    'Delivery timing could be better',
    'Returns process was simple',
    'The product broke after one week',
    'Excellent value and quick response from support'
]
more_ids = range(6, 14)
df2 = pd.DataFrame({'CustomerID': list(more_ids), 'Feedback': more_data})
df = pd.concat([df, df2], ignore_index=True)
print(df.shape)
print(df.tail())
(13, 2)
    CustomerID                                         Feedback
8            9                Always a great deal at this store
9           10                  Delivery timing could be better
10          11                       Returns process was simple
11          12                 The product broke after one week
12          13  Excellent value and quick response from support
# Create a word cloud from the larger feedback set
feedback_text = ' '.join(df['Feedback']).lower()
wc = WordCloud(width=520, height=260, background_color='white', stopwords=STOPWORDS, max_words=30).generate(feedback_text)
plt.figure(figsize=(10,5))
plt.imshow(wc, interpolation='bilinear')
plt.axis('off')
plt.title('Customer Feedback: Top 30 Words')
plt.show()
No description has been provided for this image
# Bar plot of most common feedback words
from collections import Counter
words = feedback_text.split()
filtered_words = [w for w in words if w not in stopwords]
top_words = Counter(filtered_words).most_common(10)
words, counts = zip(*top_words)
plt.figure(figsize=(8,4))
sns.barplot(x=list(counts), y=list(words), palette='Blues_r')
plt.xlabel('Count')
plt.title('Top 10 Words in Customer Feedback')
plt.tight_layout()
plt.show()
No description has been provided for this image
# Adding sentiment: quick manual tagging
def tag_sentiment(text):
    positives = ['great', 'excellent', 'good', 'fast', 'clean', 'value', 'simple', 'quick', 'help']
    negatives = ['poor', 'slow', 'broke', 'confusing', 'did not', 'needs', 'not']
    text_lower = text.lower()
    if any(word in text_lower for word in positives):
        return 'Positive'
    elif any(word in text_lower for word in negatives):
        return 'Negative'
    else:
        return 'Neutral'
df['Sentiment'] = df['Feedback'].apply(tag_sentiment)
print(df[['Feedback', 'Sentiment']].head(8))
                                   Feedback Sentiment
0          Great service and friendly staff  Positive
1  Delivery was slow and packaging was poor  Negative
2         Excellent quality, will buy again  Positive
3        Customer support needs improvement  Negative
4                      Good value for money  Positive
5    Staff did not help me with my question  Positive
6         Fast shipping and clean packaging  Positive
7             I found the website confusing  Negative
# Visualize distribution of sentiments
sent_counts = df['Sentiment'].value_counts()
sns.barplot(x=sent_counts.index, y=sent_counts.values, palette='viridis')
plt.xlabel('Sentiment')
plt.ylabel('Number of Comments')
plt.title('Feedback Sentiment Distribution')
plt.tight_layout()
plt.show()
No description has been provided for this image
# Feedback themes by sentiment: intermediate pattern
for sentiment in ['Positive','Negative','Neutral']:
    subset = df[df['Sentiment']==sentiment]
    if len(subset):
        subset_text = ' '.join(subset['Feedback']).lower()
        sub_wc = WordCloud(width=420, height=190, background_color='white', stopwords=STOPWORDS, max_words=10).generate(subset_text)
        plt.figure(figsize=(6,3))
        plt.imshow(sub_wc, interpolation='bilinear')
        plt.axis('off')
        plt.title(f'{sentiment} Feedback: Top Words')
        plt.show()
No description has been provided for this image
No description has been provided for this image
No description has been provided for this image
# Connecting text feedback to scores: setup NPS example
np.random.seed(42)
nps_scores = np.random.randint(0, 11, size=len(df))
df['NPS_Score'] = nps_scores
print(df[['Feedback','NPS_Score']].head())
                                   Feedback  NPS_Score
0          Great service and friendly staff          6
1  Delivery was slow and packaging was poor          3
2         Excellent quality, will buy again         10
3        Customer support needs improvement          7
4                      Good value for money          4
# Visualize average NPS by sentiment
nps_by_sent = df.groupby('Sentiment')['NPS_Score'].mean()
sns.barplot(x=nps_by_sent.index, y=nps_by_sent.values, palette='mako')
plt.ylabel('Average NPS Score')
plt.title('Average NPS by Feedback Sentiment')
plt.tight_layout()
plt.show()
No description has been provided for this image
# Intermediate: Find and visualize top trigrams in feedback
from sklearn.feature_extraction.text import CountVectorizer
vectorizer = CountVectorizer(stop_words='english', ngram_range=(3,3))
X = vectorizer.fit_transform(df['Feedback'])
sum_words = X.sum(axis=0)
words_freq = [(word, sum_words[0, idx]) for word, idx in vectorizer.vocabulary_.items()]
words_freq = sorted(words_freq, key=lambda x: x[1], reverse=True)[:7]
trigrams, counts = zip(*words_freq) if words_freq else ([], [])
plt.figure(figsize=(10,3))
sns.barplot(x=list(counts), y=list(trigrams), palette='flare')
plt.xlabel('Count')
plt.title('Most Common Trigrams in Feedback')
plt.tight_layout()
plt.show()
No description has been provided for this image
# Advanced: Connect text feedback to customer demographics
ages = np.random.randint(18, 65, len(df))
regions = np.random.choice(['North', 'South', 'East', 'West'], len(df))
df['Age'] = ages
df['Region'] = regions
plt.figure(figsize=(7,4))
sns.boxplot(x='Region', y='NPS_Score', data=df, palette='pastel')
plt.title('NPS Score Distribution by Region')
plt.show()
No description has been provided for this image
# Advanced: Topic separation using simple keyword matching
df['Topic'] = df['Feedback'].str.contains('delivery|shipping|packaging|timing', case=False).map({True: 'Delivery', False: 'Other'})
topic_counts = df.groupby(['Topic','Sentiment']).size().unstack(fill_value=0)
topic_counts.plot(kind='bar', stacked=True, color=['green','red','grey'], figsize=(7,4))
plt.title('Topic vs Sentiment in Feedback')
plt.ylabel('Number of Comments')
plt.xticks(rotation=0)
plt.tight_layout()
plt.show()
No description has been provided for this image
# Error handling 1: Missing comment fills
df_with_missing = df.copy()
df_with_missing.loc[2,'Feedback'] = np.nan
df_with_missing['Feedback_filled'] = df_with_missing['Feedback'].fillna('no feedback provided')
print(df_with_missing[['CustomerID','Feedback','Feedback_filled']].head(5))
   CustomerID                                  Feedback  \
0           1          Great service and friendly staff   
1           2  Delivery was slow and packaging was poor   
2           3                                       NaN   
3           4        Customer support needs improvement   
4           5                      Good value for money   

                            Feedback_filled  
0          Great service and friendly staff  
1  Delivery was slow and packaging was poor  
2                      no feedback provided  
3        Customer support needs improvement  
4                      Good value for money  
# Error handling 2: Incorrect groupings
try:
    # Wrong: Group by non-existing column
    test = df.groupby('Sent').size()
except Exception as e:
    print('Grouping error:', e)
Grouping error: 'Sent'
# Error handling 3: Misinterpreting sentiment and NPS
df['Likert_Score'] = np.where(df['NPS_Score'] >= 8, 'Agree', np.where(df['NPS_Score'] <= 3, 'Disagree', 'Neutral'))
crosstab = pd.crosstab(df['Sentiment'], df['Likert_Score'])
print(crosstab)
Likert_Score  Agree  Disagree  Neutral
Sentiment                             
Negative          0         2        2
Neutral           1         0        0
Positive          3         0        5

Best Practices: Patterns for Text Market Research#

  • Segment feedback by meaningful groups: product, location, score tier.
  • Always remove stopwords and check for spelling variations.
  • Use cross-tabs to relate text sentiment to numeric ratings.
  • Build indexes or thematic scores by counting major topics.
  • Visualize trends over time if you have dates.
# Pattern: Index construction (complaints index)
df['Is_Complaint'] = df['Feedback'].str.contains('broke|did not|needs|poor|confusing|not help', case=False)
complaint_index = df['Is_Complaint'].mean()
print(f'Complaint Index: {complaint_index:.2%} of comments are complaints')
Complaint Index: 38.46% of comments are complaints
# Pattern: Follow-up deep dive on frequent complaint
complaints = df[df['Is_Complaint']]
print('Sample complaints:')
print(complaints['Feedback'].tolist()[:5])
Sample complaints:
['Delivery was slow and packaging was poor', 'Customer support needs improvement', 'Staff did not help me with my question', 'I found the website confusing', 'The product broke after one week']

End-to-End Example: Insight to Recommendation#

  • We start with customer feedback text and basic demographic segmentation.
  • We identify dominant positive and negative topics using word clouds and bar plots.
  • Sentiment analysis and complaint indexing provide numeric metrics.
  • We connect these insights to NPS for actionable reporting.
  • Recommendations: Focus improvements on most-frequent complaint topics and regions with lowest average NPS.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.