Lesson 24 · Market Research Analytics in Python
Visualizing Survey and Market Data Using Python for Market Research
In this lesson, we will explore how to turn survey and customer responses into clear business insights using Python visualizations. Understanding and…
- CourseMarket Research Analytics in Python
- Lesson24 of 56
- Video25 min
- FormatJupyter notebook · 22 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbVisualizing Survey and Market Data in Python#
- In this lesson, we will explore how to turn survey and customer responses into clear business insights using Python visualizations.
- Understanding and communicating survey results is critical for making evidence-based decisions in marketing, customer experience, and product development.
- You will learn how to work with real customer and market research datasets, visualize satisfaction and retention, and interpret key metrics.
- By the end, you will be able to analyze and visualize survey data, spot trends, and communicate findings clearly.
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns
import openml
import warnings
warnings.filterwarnings('ignore')
Understanding Survey and Market Data#
- Survey datasets often include responses, satisfaction ratings, NPS scores, and customer demographics such as age, gender, and location.
- Customer analytics data can also include purchasing behaviors, churn, and feedback text.
- Each row typically represents a customer or respondent.
- Common mistakes include treating categorical values as continuous, ignoring missing data, or failing to segment by demographic factors.
- Always double-check what each column means before visualizing. Proper structure leads to meaningful charts.
# Load customer satisfaction survey data from OpenML
dataset = openml.datasets.get_dataset(42178)
df_satis, _, _, _ = dataset.get_data(dataset_format='dataframe')
print(df_satis.shape)
print(df_satis.head(3))
# Load marketing campaign performance data
dataset = openml.datasets.get_dataset(1461)
df_marketing, _, _, _ = dataset.get_data(dataset_format='dataframe')
df_marketing.columns = ['age','job','marital','education','default','balance','housing','loan','contact','day','month','duration','campaign','pdays','previous','poutcome','response']
print(df_marketing.shape)
print(df_marketing.head(3))
# Load synthetic Net Promoter Score (NPS) survey data
np.random.seed(42)
df_nps = pd.DataFrame({'CustomerID': range(1,501),
'Age': np.random.randint(18,70,500),
'Region': np.random.choice(['North','South','East','West'],500),
'NPS_Score': np.random.randint(0,11,500)})
print(df_nps.shape)
print(df_nps.head(3))
# Beginner Example 1: Count survey responses by gender
ax = df_satis['gender'].value_counts().plot(kind='bar', color='skyblue')
plt.title('Number of Survey Respondents by Gender')
plt.xlabel('Gender')
plt.ylabel('Count')
plt.tight_layout()
plt.show()
# Beginner Example 2: Plot NPS Score distribution
plt.figure(figsize=(8,4))
sns.histplot(df_nps['NPS_Score'], bins=11, kde=False, color='green')
plt.title('Distribution of NPS Scores')
plt.xlabel('NPS Score')
plt.ylabel('Customer Count')
plt.tight_layout()
plt.show()
# Beginner Example 3: Pie chart of marketing campaign responses
response_counts = df_marketing['response'].value_counts()
plt.figure(figsize=(6,6))
plt.pie(response_counts, labels=response_counts.index, autopct='%1.1f%%', colors=['#89cff0','#ffcccb'])
plt.title('Customer Response Rate in Campaign')
plt.show()
# Intermediate Example 1: NPS by region
region_nps = df_nps.groupby('Region')['NPS_Score'].mean()
region_nps.plot(kind='bar', color=['#0072B2', '#D55E00', '#CC79A7', '#009E73'])
plt.title('Average NPS Score by Region')
plt.xlabel('Region')
plt.ylabel('Average NPS Score')
plt.ylim(0,10)
plt.tight_layout()
plt.show()
# Intermediate Example 2: Monthly charges vs churn rate
churn_rate = df_satis.groupby('MonthlyCharges')['Churn'].apply(lambda x: (x == 'Yes').mean())
plt.figure(figsize=(10,4))
plt.plot(churn_rate.index, churn_rate.values, marker='o', linestyle='-')
plt.title('Churn Rate by Monthly Charges')
plt.xlabel('Monthly Charges')
plt.ylabel('Churn Rate')
plt.grid(True)
plt.tight_layout()
plt.show()
# Intermediate Example 3: Cross-tabulation heatmap of Internet service vs Churn
crosstab = pd.crosstab(df_satis['InternetService'], df_satis['Churn'])
sns.heatmap(crosstab, annot=True, fmt='d', cmap='YlGnBu')
plt.title('Churn by Internet Service Type')
plt.xlabel('Churn')
plt.ylabel('Internet Service')
plt.tight_layout()
plt.show()
# Advanced Example 1: NPS segmentation by age group and region
df_nps['AgeGroup'] = pd.cut(df_nps['Age'], bins=[17,29,39,49,59,70], labels=['18-29','30-39','40-49','50-59','60-69'])
pivot = df_nps.pivot_table(index='AgeGroup', columns='Region', values='NPS_Score', aggfunc='mean')
plt.figure(figsize=(8,5))
sns.heatmap(pivot, annot=True, cmap='coolwarm', center=5, linewidths=0.5)
plt.title('Average NPS Score by Age Group and Region')
plt.ylabel('Age Group')
plt.xlabel('Region')
plt.tight_layout()
plt.show()
# Advanced Example 2: Visualizing customer retention trends using cohort analysis
np.random.seed(0)
dates = pd.date_range('2021-01-01', periods=24, freq='ME')
df_cohort = pd.DataFrame({'CustomerID': np.random.randint(1000,2000,len(dates)),
'Signup_Month': dates,
'Active_Users': np.random.randint(50,300,len(dates))})
plt.figure(figsize=(10,4))
plt.plot(df_cohort['Signup_Month'], df_cohort['Active_Users'], marker='o')
plt.title('Customer Retention Trend Over Time')
plt.xlabel('Signup Month')
plt.ylabel('Active Users')
plt.xticks(rotation=45)
plt.tight_layout()
plt.show()
# Advanced Example 3: Word cloud of open-ended customer feedback
from wordcloud import WordCloud, STOPWORDS
df_feedback = pd.DataFrame({'CustomerID':[1,2,3,4,5],
'Feedback':['Great service and friendly staff','Delivery was slow and packaging was poor',
'Excellent quality, will buy again','Customer support needs improvement','Good value for money']})
text = ' '.join(df_feedback['Feedback'])
wordcloud = WordCloud(stopwords=STOPWORDS, background_color='white', width=600, height=400).generate(text)
plt.figure(figsize=(8,4))
plt.imshow(wordcloud, interpolation='bilinear')
plt.axis('off')
plt.title('Customer Feedback Highlights')
plt.show()
# Error Handling Example 1: Detect missing values in survey data
missing_counts = df_satis.isnull().sum()
print('Missing values per column:')
print(missing_counts[missing_counts > 0])
# Error Handling Example 2: Incorrect groupby on categorical fields
try:
result = df_marketing.groupby('balance')['age'].mean()
print(result.head())
except Exception as e:
print('Error:', e)
print('Check if grouping field is too granular and should be binned.')
# Error Handling Example 3: Misinterpreting NPS score distribution
promoters = (df_nps['NPS_Score'] >= 9).sum()
detractors = (df_nps['NPS_Score'] <= 6).sum()
passives = ((df_nps['NPS_Score'] >= 7) & (df_nps['NPS_Score'] <= 8)).sum()
total = len(df_nps)
nps = ((promoters - detractors) / total) * 100
print(f'NPS = {nps:.1f} (should range from -100 to 100)')
Best Practices in Market Research Analytics#
- Always segment data by relevant business characteristics such as age, region, or service tier.
- Use cross-tabulations and heatmaps to detect patterns and outliers quickly.
- Build indices and scores using official methods (such as NPS) and explain ranges to business users.
- Monitor trends over time to capture changes in satisfaction, campaign, or retention metrics.
- Double-check visualizations for missing or misclassified data before presenting findings.
# Pattern: Customer segmentation by satisfaction level
df_satis['Satisfaction_Level'] = pd.cut(df_satis['MonthlyCharges'], bins=[0,30,60,90,120], labels=['Basic','Standard','Premium','Elite'])
segment_counts = df_satis['Satisfaction_Level'].value_counts().sort_index()
segment_counts.plot(kind='bar', color='orange')
plt.title('Customer Segmentation by Monthly Charges')
plt.xlabel('Segment')
plt.ylabel('Customer Count')
plt.tight_layout()
plt.show()
# Pattern: Cross-tab of churn by payment method
crosstab2 = pd.crosstab(df_satis['PaymentMethod'], df_satis['Churn'], normalize='index')
plt.figure(figsize=(8,5))
sns.heatmap(crosstab2, annot=True, cmap='viridis', fmt='.2f')
plt.title('Churn Rate by Payment Method (Proportion)')
plt.xlabel('Churn')
plt.ylabel('Payment Method')
plt.tight_layout()
plt.show()
# Pattern: Trend analysis using moving average on active users
df_cohort['MA_3'] = df_cohort['Active_Users'].rolling(window=3, min_periods=1).mean()
plt.figure(figsize=(10,4))
plt.plot(df_cohort['Signup_Month'], df_cohort['Active_Users'], label='Active Users')
plt.plot(df_cohort['Signup_Month'], df_cohort['MA_3'], label='3-Month Moving Average', linestyle='--')
plt.title('Customer Retention Trend - 3 Month Smoothing')
plt.xlabel('Signup Month')
plt.ylabel('Active Users')
plt.legend()
plt.tight_layout()
plt.show()
# End-to-End Example: From survey data to business recommendation
# Step 1: Visualize satisfaction by contract type
contract_churn = pd.crosstab(df_satis['Contract'], df_satis['Churn'], normalize='index')
contract_churn.plot(kind='bar', stacked=True, color=['#e74c3c','#2ecc71'])
plt.title('Churn Rate by Contract Type')
plt.xlabel('Contract Type')
plt.ylabel('Proportion')
plt.legend(title='Churn')
plt.tight_layout()
plt.show()
# Step 2: Recommend a churn reduction action
highest_churn = contract_churn['Yes'].idxmax()
print(f'Customers on {highest_churn} contracts have the highest churn. Consider retention offers targeting this group.')
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



