Mathew K Analytics

Lesson 24 · Market Research Analytics in Python

Visualizing Survey and Market Data Using Python for Market Research

In this lesson, we will explore how to turn survey and customer responses into clear business insights using Python visualizations. Understanding and…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Visualizing Survey and Market Data in Python#

  • In this lesson, we will explore how to turn survey and customer responses into clear business insights using Python visualizations.
  • Understanding and communicating survey results is critical for making evidence-based decisions in marketing, customer experience, and product development.
  • You will learn how to work with real customer and market research datasets, visualize satisfaction and retention, and interpret key metrics.
  • By the end, you will be able to analyze and visualize survey data, spot trends, and communicate findings clearly.
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns
import openml
import warnings
warnings.filterwarnings('ignore')

Understanding Survey and Market Data#

  • Survey datasets often include responses, satisfaction ratings, NPS scores, and customer demographics such as age, gender, and location.
  • Customer analytics data can also include purchasing behaviors, churn, and feedback text.
  • Each row typically represents a customer or respondent.
  • Common mistakes include treating categorical values as continuous, ignoring missing data, or failing to segment by demographic factors.
  • Always double-check what each column means before visualizing. Proper structure leads to meaningful charts.
# Load customer satisfaction survey data from OpenML
dataset = openml.datasets.get_dataset(42178)
df_satis, _, _, _ = dataset.get_data(dataset_format='dataframe')
print(df_satis.shape)
print(df_satis.head(3))
(7043, 20)
   gender  SeniorCitizen Partner Dependents  tenure PhoneService  \
0  Female              0     Yes         No       1           No   
1    Male              0      No         No      34          Yes   
2    Male              0      No         No       2          Yes   

      MultipleLines InternetService OnlineSecurity OnlineBackup  \
0  No phone service             DSL             No          Yes   
1                No             DSL            Yes           No   
2                No             DSL            Yes          Yes   

  DeviceProtection TechSupport StreamingTV StreamingMovies        Contract  \
0               No          No          No              No  Month-to-month   
1              Yes          No          No              No        One year   
2               No          No          No              No  Month-to-month   

  PaperlessBilling     PaymentMethod  MonthlyCharges TotalCharges Churn  
0              Yes  Electronic check           29.85        29.85    No  
1               No      Mailed check           56.95       1889.5    No  
2              Yes      Mailed check           53.85       108.15   Yes  
# Load marketing campaign performance data
dataset = openml.datasets.get_dataset(1461)
df_marketing, _, _, _ = dataset.get_data(dataset_format='dataframe')
df_marketing.columns = ['age','job','marital','education','default','balance','housing','loan','contact','day','month','duration','campaign','pdays','previous','poutcome','response']
print(df_marketing.shape)
print(df_marketing.head(3))
(45211, 17)
   age           job  marital  education default  balance housing loan  \
0   58    management  married   tertiary      no   2143.0     yes   no   
1   44    technician   single  secondary      no     29.0     yes   no   
2   33  entrepreneur  married  secondary      no      2.0     yes  yes   

   contact  day month  duration  campaign  pdays  previous poutcome response  
0  unknown    5   may     261.0         1   -1.0       0.0  unknown        1  
1  unknown    5   may     151.0         1   -1.0       0.0  unknown        1  
2  unknown    5   may      76.0         1   -1.0       0.0  unknown        1  
# Load synthetic Net Promoter Score (NPS) survey data
np.random.seed(42)
df_nps = pd.DataFrame({'CustomerID': range(1,501),
                      'Age': np.random.randint(18,70,500),
                      'Region': np.random.choice(['North','South','East','West'],500),
                      'NPS_Score': np.random.randint(0,11,500)})
print(df_nps.shape)
print(df_nps.head(3))
(500, 4)
   CustomerID  Age Region  NPS_Score
0           1   56   West          2
1           2   69  North          0
2           3   46   East          4
# Beginner Example 1: Count survey responses by gender
ax = df_satis['gender'].value_counts().plot(kind='bar', color='skyblue')
plt.title('Number of Survey Respondents by Gender')
plt.xlabel('Gender')
plt.ylabel('Count')
plt.tight_layout()
plt.show()
No description has been provided for this image
# Beginner Example 2: Plot NPS Score distribution
plt.figure(figsize=(8,4))
sns.histplot(df_nps['NPS_Score'], bins=11, kde=False, color='green')
plt.title('Distribution of NPS Scores')
plt.xlabel('NPS Score')
plt.ylabel('Customer Count')
plt.tight_layout()
plt.show()
No description has been provided for this image
# Beginner Example 3: Pie chart of marketing campaign responses
response_counts = df_marketing['response'].value_counts()
plt.figure(figsize=(6,6))
plt.pie(response_counts, labels=response_counts.index, autopct='%1.1f%%', colors=['#89cff0','#ffcccb'])
plt.title('Customer Response Rate in Campaign')
plt.show()
No description has been provided for this image
# Intermediate Example 1: NPS by region
region_nps = df_nps.groupby('Region')['NPS_Score'].mean()
region_nps.plot(kind='bar', color=['#0072B2', '#D55E00', '#CC79A7', '#009E73'])
plt.title('Average NPS Score by Region')
plt.xlabel('Region')
plt.ylabel('Average NPS Score')
plt.ylim(0,10)
plt.tight_layout()
plt.show()
No description has been provided for this image
# Intermediate Example 2: Monthly charges vs churn rate
churn_rate = df_satis.groupby('MonthlyCharges')['Churn'].apply(lambda x: (x == 'Yes').mean())
plt.figure(figsize=(10,4))
plt.plot(churn_rate.index, churn_rate.values, marker='o', linestyle='-')
plt.title('Churn Rate by Monthly Charges')
plt.xlabel('Monthly Charges')
plt.ylabel('Churn Rate')
plt.grid(True)
plt.tight_layout()
plt.show()
No description has been provided for this image
# Intermediate Example 3: Cross-tabulation heatmap of Internet service vs Churn
crosstab = pd.crosstab(df_satis['InternetService'], df_satis['Churn'])
sns.heatmap(crosstab, annot=True, fmt='d', cmap='YlGnBu')
plt.title('Churn by Internet Service Type')
plt.xlabel('Churn')
plt.ylabel('Internet Service')
plt.tight_layout()
plt.show()
No description has been provided for this image
# Advanced Example 1: NPS segmentation by age group and region
df_nps['AgeGroup'] = pd.cut(df_nps['Age'], bins=[17,29,39,49,59,70], labels=['18-29','30-39','40-49','50-59','60-69'])
pivot = df_nps.pivot_table(index='AgeGroup', columns='Region', values='NPS_Score', aggfunc='mean')
plt.figure(figsize=(8,5))
sns.heatmap(pivot, annot=True, cmap='coolwarm', center=5, linewidths=0.5)
plt.title('Average NPS Score by Age Group and Region')
plt.ylabel('Age Group')
plt.xlabel('Region')
plt.tight_layout()
plt.show()
No description has been provided for this image
# Advanced Example 2: Visualizing customer retention trends using cohort analysis
np.random.seed(0)
dates = pd.date_range('2021-01-01', periods=24, freq='ME')
df_cohort = pd.DataFrame({'CustomerID': np.random.randint(1000,2000,len(dates)),
                         'Signup_Month': dates,
                         'Active_Users': np.random.randint(50,300,len(dates))})
plt.figure(figsize=(10,4))
plt.plot(df_cohort['Signup_Month'], df_cohort['Active_Users'], marker='o')
plt.title('Customer Retention Trend Over Time')
plt.xlabel('Signup Month')
plt.ylabel('Active Users')
plt.xticks(rotation=45)
plt.tight_layout()
plt.show()
No description has been provided for this image
# Advanced Example 3: Word cloud of open-ended customer feedback
from wordcloud import WordCloud, STOPWORDS
df_feedback = pd.DataFrame({'CustomerID':[1,2,3,4,5],
    'Feedback':['Great service and friendly staff','Delivery was slow and packaging was poor',
    'Excellent quality, will buy again','Customer support needs improvement','Good value for money']})
text = ' '.join(df_feedback['Feedback'])
wordcloud = WordCloud(stopwords=STOPWORDS, background_color='white', width=600, height=400).generate(text)
plt.figure(figsize=(8,4))
plt.imshow(wordcloud, interpolation='bilinear')
plt.axis('off')
plt.title('Customer Feedback Highlights')
plt.show()
No description has been provided for this image
# Error Handling Example 1: Detect missing values in survey data
missing_counts = df_satis.isnull().sum()
print('Missing values per column:')
print(missing_counts[missing_counts > 0])
Missing values per column:
Series([], dtype: int64)
# Error Handling Example 2: Incorrect groupby on categorical fields
try:
    result = df_marketing.groupby('balance')['age'].mean()
    print(result.head())
except Exception as e:
    print('Error:', e)
    print('Check if grouping field is too granular and should be binned.')
balance
-8019.0    26.0
-6847.0    49.0
-4057.0    60.0
-3372.0    43.0
-3313.0    57.0
Name: age, dtype: float64
# Error Handling Example 3: Misinterpreting NPS score distribution
promoters = (df_nps['NPS_Score'] >= 9).sum()
detractors = (df_nps['NPS_Score'] <= 6).sum()
passives = ((df_nps['NPS_Score'] >= 7) & (df_nps['NPS_Score'] <= 8)).sum()
total = len(df_nps)
nps = ((promoters - detractors) / total) * 100
print(f'NPS = {nps:.1f} (should range from -100 to 100)')
NPS = -47.6 (should range from -100 to 100)

Best Practices in Market Research Analytics#

  • Always segment data by relevant business characteristics such as age, region, or service tier.
  • Use cross-tabulations and heatmaps to detect patterns and outliers quickly.
  • Build indices and scores using official methods (such as NPS) and explain ranges to business users.
  • Monitor trends over time to capture changes in satisfaction, campaign, or retention metrics.
  • Double-check visualizations for missing or misclassified data before presenting findings.
# Pattern: Customer segmentation by satisfaction level
df_satis['Satisfaction_Level'] = pd.cut(df_satis['MonthlyCharges'], bins=[0,30,60,90,120], labels=['Basic','Standard','Premium','Elite'])
segment_counts = df_satis['Satisfaction_Level'].value_counts().sort_index()
segment_counts.plot(kind='bar', color='orange')
plt.title('Customer Segmentation by Monthly Charges')
plt.xlabel('Segment')
plt.ylabel('Customer Count')
plt.tight_layout()
plt.show()
No description has been provided for this image
# Pattern: Cross-tab of churn by payment method
crosstab2 = pd.crosstab(df_satis['PaymentMethod'], df_satis['Churn'], normalize='index')
plt.figure(figsize=(8,5))
sns.heatmap(crosstab2, annot=True, cmap='viridis', fmt='.2f')
plt.title('Churn Rate by Payment Method (Proportion)')
plt.xlabel('Churn')
plt.ylabel('Payment Method')
plt.tight_layout()
plt.show()
No description has been provided for this image
# Pattern: Trend analysis using moving average on active users
df_cohort['MA_3'] = df_cohort['Active_Users'].rolling(window=3, min_periods=1).mean()
plt.figure(figsize=(10,4))
plt.plot(df_cohort['Signup_Month'], df_cohort['Active_Users'], label='Active Users')
plt.plot(df_cohort['Signup_Month'], df_cohort['MA_3'], label='3-Month Moving Average', linestyle='--')
plt.title('Customer Retention Trend - 3 Month Smoothing')
plt.xlabel('Signup Month')
plt.ylabel('Active Users')
plt.legend()
plt.tight_layout()
plt.show()
No description has been provided for this image
# End-to-End Example: From survey data to business recommendation
# Step 1: Visualize satisfaction by contract type
contract_churn = pd.crosstab(df_satis['Contract'], df_satis['Churn'], normalize='index')
contract_churn.plot(kind='bar', stacked=True, color=['#e74c3c','#2ecc71'])
plt.title('Churn Rate by Contract Type')
plt.xlabel('Contract Type')
plt.ylabel('Proportion')
plt.legend(title='Churn')
plt.tight_layout()
plt.show()
No description has been provided for this image
# Step 2: Recommend a churn reduction action
highest_churn = contract_churn['Yes'].idxmax()
print(f'Customers on {highest_churn} contracts have the highest churn. Consider retention offers targeting this group.')
Customers on Month-to-month contracts have the highest churn. Consider retention offers targeting this group.
 

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.