Lesson 54 · Market Research Analytics in Python
Best Practices for Market Research Visuals | Python Analytics Tutorial
Visualizing market research data is important for understanding customer needs, trends, and opportunities. Market research visuals help business leaders…
- CourseMarket Research Analytics in Python
- Lesson54 of 56
- Video27 min
- FormatJupyter notebook · 27 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbBest Practices for Market Research Visuals#
- Visualizing market research data is important for understanding customer needs, trends, and opportunities.
- Market research visuals help business leaders make informed decisions and communicate insights clearly.
- In this lesson, you will explore data, create effective charts, fix common visualization mistakes, and design clear market research dashboards.
- By the end, you will produce actionable insights and visualizations from real survey and customer data.
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns
import openml
import warnings
warnings.filterwarnings('ignore')
Core market research concepts for visuals#
- Surveys capture demographics, attitudes, satisfaction ratings, and text feedback.
- Customer data may be structured (numbers, categories) or unstructured (text comments).
- Charts must match data type (bar for counts, boxplot for scores, wordcloud for text).
- Beginners often use the wrong chart, skip labels, or mislead with axes.
# Load customer satisfaction dataset (OpenML 42178)
dataset = openml.datasets.get_dataset(42178)
df, _, _, _ = dataset.get_data(dataset_format='dataframe')
print('Shape:', df.shape)
print(df.head(3))
# Simple bar chart: Gender distribution
gender_counts = df['gender'].value_counts()
plt.figure(figsize=(5,3))
sns.barplot(x=gender_counts.index, y=gender_counts.values, palette='pastel')
plt.title('Customer Gender Distribution')
plt.ylabel('Number of Customers')
plt.xlabel('Gender')
plt.tight_layout()
plt.show()
# Pie chart: Customer Churn Rate
churn_counts = df['Churn'].value_counts()
plt.figure(figsize=(4,4))
plt.pie(churn_counts, labels=churn_counts.index, autopct='%1.1f%%', colors=['#5ab4ac','#d8b365'])
plt.title('Customer Churn Rate')
plt.show()
# Countplot: Customer Contract Types
plt.figure(figsize=(6,3))
sns.countplot(x='Contract', data=df, palette='Set2')
plt.title('Distribution of Contract Types')
plt.xlabel('Contract Type')
plt.ylabel('Number of Customers')
plt.tight_layout()
plt.show()
# Histogram: Monthly Charges Distribution
plt.figure(figsize=(6,3))
plt.hist(df['MonthlyCharges'], bins=20, color='#4575b4', edgecolor='black')
plt.title('Monthly Charges Distribution')
plt.xlabel('Monthly Charges ($)')
plt.ylabel('Number of Customers')
plt.tight_layout()
plt.show()
# Boxplot: Monthly Charges by Churn Status
plt.figure(figsize=(6,4))
sns.boxplot(x='Churn', y='MonthlyCharges', data=df, palette=['#99d594','#fc8d62'])
plt.title('Monthly Charges by Churn Status')
plt.xlabel('Churn')
plt.ylabel('Monthly Charges ($)')
plt.tight_layout()
plt.show()
# Stacked Bar Chart: Internet Service Type by Churn
ct = pd.crosstab(df['InternetService'], df['Churn'])
ct.plot(kind='bar', stacked=True, color=['#5ab4ac','#d8b365'], figsize=(7,4))
plt.title('Internet Service Type by Churn')
plt.xlabel('Internet Service Type')
plt.ylabel('Number of Customers')
plt.xticks(rotation=0)
plt.legend(title='Churn')
plt.tight_layout()
plt.show()
# Trend Line: Average Monthly Charges by Tenure
avg_charges_by_tenure = df.groupby('tenure')['MonthlyCharges'].mean()
plt.figure(figsize=(8,3))
plt.plot(avg_charges_by_tenure.index, avg_charges_by_tenure.values, color='#fee08b', marker='o')
plt.title('Average Monthly Charges by Tenure (Months)')
plt.xlabel('Tenure (Months)')
plt.ylabel('Avg Monthly Charges')
plt.grid(True, linestyle=':', alpha=0.5)
plt.tight_layout()
plt.show()
# Setup: Generate synthetic NPS survey dataset
np.random.seed(42)
nps_df = pd.DataFrame({
'CustomerID': range(1,501),
'Age': np.random.randint(18,70,500),
'Region': np.random.choice(['North','South','East','West'],500),
'NPS_Score': np.random.randint(0,11,500)
})
print(nps_df.head(3))
# Bar chart: NPS Score Distribution
plt.figure(figsize=(7,3))
sns.countplot(x='NPS_Score', data=nps_df, palette='Blues')
plt.title('Customer NPS Score Distribution')
plt.xlabel('NPS Score (0=Detractor, 10=Promoter)')
plt.ylabel('Number of Responses')
plt.tight_layout()
plt.show()
# Grouped Bar: Mean NPS Score by Region
nps_region = nps_df.groupby('Region')['NPS_Score'].mean().reset_index()
plt.figure(figsize=(6,3))
sns.barplot(x='Region', y='NPS_Score', data=nps_region, palette='Set1')
plt.title('Average NPS Score by Region')
plt.ylabel('Average NPS Score')
plt.xlabel('Region')
plt.tight_layout()
plt.show()
# Setup: Sample open-ended customer feedback
feedback_df = pd.DataFrame({
'CustomerID':[1,2,3,4,5],
'Feedback':[
'Great service and friendly staff',
'Delivery was slow and packaging was poor',
'Excellent quality, will buy again',
'Customer support needs improvement',
'Good value for money']
})
print(feedback_df)
# Wordcloud for customer feedback
from wordcloud import WordCloud
text = ' '.join(feedback_df['Feedback'])
wordcloud = WordCloud(background_color='white', width=400, height=200).generate(text)
plt.figure(figsize=(7,3))
plt.imshow(wordcloud, interpolation='bilinear')
plt.axis('off')
plt.title('Customer Feedback Wordcloud')
plt.show()
# Advanced: Cluster customers by tenure and monthly charges
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
clust_df = df[['tenure','MonthlyCharges']].dropna()
scaler = StandardScaler()
scaled = scaler.fit_transform(clust_df)
kmeans = KMeans(n_clusters=3, random_state=42)
labels = kmeans.fit_predict(scaled)
clust_df['Cluster'] = labels
plt.figure(figsize=(7,4))
sns.scatterplot(x='tenure', y='MonthlyCharges', hue='Cluster', data=clust_df, palette='viridis')
plt.title('Market Segmentation by Tenure & Charges')
plt.xlabel('Tenure (Months)')
plt.ylabel('Monthly Charges ($)')
plt.legend(title='Cluster')
plt.tight_layout()
plt.show()
# Generate synthetic cohort dataset
np.random.seed(0)
dates = pd.date_range('2021-01-01', periods=24, freq='ME')
cohort_df = pd.DataFrame({
'CustomerID': np.random.randint(1000,2000,len(dates)),
'Signup_Month': dates,
'Active_Users': np.random.randint(50,300,len(dates))
})
print(cohort_df.head(3))
# Line plot: Cohort Retention Trend
plt.figure(figsize=(8,4))
plt.plot(cohort_df['Signup_Month'], cohort_df['Active_Users'], marker='s', color='#b2182b')
plt.title('Monthly Active Users by Cohort Signup Month')
plt.xlabel('Signup Month')
plt.ylabel('Active Users')
plt.grid(True, ls=':', alpha=0.4)
plt.tight_layout()
plt.show()
# Detect missing values in survey data
missing = df.isnull().sum()
print('Missing survey responses per column:')
print(missing[missing > 0])
# Visualize missing survey data
import missingno as msno
msno.matrix(df.sample(500, random_state=42))
plt.title('Missing Survey Data Visualization')
plt.show()
# Example: Incorrect aggregation warning
try:
avg = df.groupby('Churn')['MonthlyCharges'].sum()['Yes'] / df['Churn'].value_counts()['Yes']
print('Average monthly charge (churned):', avg)
except Exception as e:
print('Aggregation error:', e)
# Example: NPS scale misinterpretation
promoters = (nps_df['NPS_Score'] >= 9).sum()
detractors = (nps_df['NPS_Score'] <= 6).sum()
nps_score = 100 * (promoters - detractors) / len(nps_df)
print('Net Promoter Score (NPS):', round(nps_score, 1))
# Best Practice: Cross-tabulation of Churn by Contract Type
ctab = pd.crosstab(df['Contract'], df['Churn'])
print(ctab)
# Build a satisfaction score index from available indicators
df['Satisfaction_Score'] = (df['tenure']/df['tenure'].max())*0.4 + (df['MonthlyCharges']/df['MonthlyCharges'].max())*0.6
print(df[['tenure','MonthlyCharges','Satisfaction_Score']].head(3))
# Setup: Load bank marketing campaign dataset
dataset2 = openml.datasets.get_dataset(1461)
df2, _, _, _ = dataset2.get_data(dataset_format='dataframe')
df2.columns = ['age','job','marital','education','default','balance','housing','loan','contact','day','month','duration','campaign','pdays','previous','poutcome','response']
month_summary = df2['month'].value_counts().sort_index()
plt.figure(figsize=(7,3))
sns.lineplot(x=month_summary.index, y=month_summary.values, marker='o', color='#5e4fa2')
plt.title('Campaign Calls by Month')
plt.xlabel('Month')
plt.ylabel('Number of Calls')
plt.grid(ls=':', alpha=0.4)
plt.tight_layout()
plt.show()
# Mini Challenge: Identify if senior citizens are more likely to churn
senior_table = pd.crosstab(df['SeniorCitizen'], df['Churn'])
print('Senior Citizen by Churn Table:')
print(senior_table)
senior_churn_rate = senior_table.loc[1,'Yes'] / senior_table.loc[1].sum()
non_senior_churn_rate = senior_table.loc[0,'Yes'] / senior_table.loc[0].sum()
print('Churn Rate (Senior):', round(senior_churn_rate*100,1),'%')
print('Churn Rate (Non-Senior):', round(non_senior_churn_rate*100,1),'%')
# Save current figure as a PDF for business reporting
fig, ax = plt.subplots(figsize=(6,3))
sns.countplot(x='Contract', data=df, palette='Set2', ax=ax)
ax.set_title('Distribution of Contract Types')
plt.tight_layout()
fig.savefig('contract_types_distribution.pdf')
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



