Mathew K Analytics

Lesson 54 · Market Research Analytics in Python

Best Practices for Market Research Visuals | Python Analytics Tutorial

Visualizing market research data is important for understanding customer needs, trends, and opportunities. Market research visuals help business leaders…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Best Practices for Market Research Visuals#

  • Visualizing market research data is important for understanding customer needs, trends, and opportunities.
  • Market research visuals help business leaders make informed decisions and communicate insights clearly.
  • In this lesson, you will explore data, create effective charts, fix common visualization mistakes, and design clear market research dashboards.
  • By the end, you will produce actionable insights and visualizations from real survey and customer data.
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns
import openml
import warnings
warnings.filterwarnings('ignore')

Core market research concepts for visuals#

  • Surveys capture demographics, attitudes, satisfaction ratings, and text feedback.
  • Customer data may be structured (numbers, categories) or unstructured (text comments).
  • Charts must match data type (bar for counts, boxplot for scores, wordcloud for text).
  • Beginners often use the wrong chart, skip labels, or mislead with axes.
# Load customer satisfaction dataset (OpenML 42178)
dataset = openml.datasets.get_dataset(42178)
df, _, _, _ = dataset.get_data(dataset_format='dataframe')
print('Shape:', df.shape)
print(df.head(3))
Shape: (7043, 20)
   gender  SeniorCitizen Partner Dependents  tenure PhoneService  \
0  Female              0     Yes         No       1           No   
1    Male              0      No         No      34          Yes   
2    Male              0      No         No       2          Yes   

      MultipleLines InternetService OnlineSecurity OnlineBackup  \
0  No phone service             DSL             No          Yes   
1                No             DSL            Yes           No   
2                No             DSL            Yes          Yes   

  DeviceProtection TechSupport StreamingTV StreamingMovies        Contract  \
0               No          No          No              No  Month-to-month   
1              Yes          No          No              No        One year   
2               No          No          No              No  Month-to-month   

  PaperlessBilling     PaymentMethod  MonthlyCharges TotalCharges Churn  
0              Yes  Electronic check           29.85        29.85    No  
1               No      Mailed check           56.95       1889.5    No  
2              Yes      Mailed check           53.85       108.15   Yes  
# Simple bar chart: Gender distribution
gender_counts = df['gender'].value_counts()
plt.figure(figsize=(5,3))
sns.barplot(x=gender_counts.index, y=gender_counts.values, palette='pastel')
plt.title('Customer Gender Distribution')
plt.ylabel('Number of Customers')
plt.xlabel('Gender')
plt.tight_layout()
plt.show()
No description has been provided for this image
# Pie chart: Customer Churn Rate
churn_counts = df['Churn'].value_counts()
plt.figure(figsize=(4,4))
plt.pie(churn_counts, labels=churn_counts.index, autopct='%1.1f%%', colors=['#5ab4ac','#d8b365'])
plt.title('Customer Churn Rate')
plt.show()
No description has been provided for this image
# Countplot: Customer Contract Types
plt.figure(figsize=(6,3))
sns.countplot(x='Contract', data=df, palette='Set2')
plt.title('Distribution of Contract Types')
plt.xlabel('Contract Type')
plt.ylabel('Number of Customers')
plt.tight_layout()
plt.show()
No description has been provided for this image
# Histogram: Monthly Charges Distribution
plt.figure(figsize=(6,3))
plt.hist(df['MonthlyCharges'], bins=20, color='#4575b4', edgecolor='black')
plt.title('Monthly Charges Distribution')
plt.xlabel('Monthly Charges ($)')
plt.ylabel('Number of Customers')
plt.tight_layout()
plt.show()
No description has been provided for this image
# Boxplot: Monthly Charges by Churn Status
plt.figure(figsize=(6,4))
sns.boxplot(x='Churn', y='MonthlyCharges', data=df, palette=['#99d594','#fc8d62'])
plt.title('Monthly Charges by Churn Status')
plt.xlabel('Churn')
plt.ylabel('Monthly Charges ($)')
plt.tight_layout()
plt.show()
No description has been provided for this image
# Stacked Bar Chart: Internet Service Type by Churn
ct = pd.crosstab(df['InternetService'], df['Churn'])
ct.plot(kind='bar', stacked=True, color=['#5ab4ac','#d8b365'], figsize=(7,4))
plt.title('Internet Service Type by Churn')
plt.xlabel('Internet Service Type')
plt.ylabel('Number of Customers')
plt.xticks(rotation=0)
plt.legend(title='Churn')
plt.tight_layout()
plt.show()
No description has been provided for this image
# Trend Line: Average Monthly Charges by Tenure
avg_charges_by_tenure = df.groupby('tenure')['MonthlyCharges'].mean()
plt.figure(figsize=(8,3))
plt.plot(avg_charges_by_tenure.index, avg_charges_by_tenure.values, color='#fee08b', marker='o')
plt.title('Average Monthly Charges by Tenure (Months)')
plt.xlabel('Tenure (Months)')
plt.ylabel('Avg Monthly Charges')
plt.grid(True, linestyle=':', alpha=0.5)
plt.tight_layout()
plt.show()
No description has been provided for this image
# Setup: Generate synthetic NPS survey dataset
np.random.seed(42)
nps_df = pd.DataFrame({
    'CustomerID': range(1,501),
    'Age': np.random.randint(18,70,500),
    'Region': np.random.choice(['North','South','East','West'],500),
    'NPS_Score': np.random.randint(0,11,500)
})
print(nps_df.head(3))
   CustomerID  Age Region  NPS_Score
0           1   56   West          2
1           2   69  North          0
2           3   46   East          4
# Bar chart: NPS Score Distribution
plt.figure(figsize=(7,3))
sns.countplot(x='NPS_Score', data=nps_df, palette='Blues')
plt.title('Customer NPS Score Distribution')
plt.xlabel('NPS Score (0=Detractor, 10=Promoter)')
plt.ylabel('Number of Responses')
plt.tight_layout()
plt.show()
No description has been provided for this image
# Grouped Bar: Mean NPS Score by Region
nps_region = nps_df.groupby('Region')['NPS_Score'].mean().reset_index()
plt.figure(figsize=(6,3))
sns.barplot(x='Region', y='NPS_Score', data=nps_region, palette='Set1')
plt.title('Average NPS Score by Region')
plt.ylabel('Average NPS Score')
plt.xlabel('Region')
plt.tight_layout()
plt.show()
No description has been provided for this image
# Setup: Sample open-ended customer feedback
feedback_df = pd.DataFrame({
    'CustomerID':[1,2,3,4,5],
    'Feedback':[
        'Great service and friendly staff',
        'Delivery was slow and packaging was poor',
        'Excellent quality, will buy again',
        'Customer support needs improvement',
        'Good value for money']
})
print(feedback_df)
   CustomerID                                  Feedback
0           1          Great service and friendly staff
1           2  Delivery was slow and packaging was poor
2           3         Excellent quality, will buy again
3           4        Customer support needs improvement
4           5                      Good value for money
# Wordcloud for customer feedback
from wordcloud import WordCloud
text = ' '.join(feedback_df['Feedback'])
wordcloud = WordCloud(background_color='white', width=400, height=200).generate(text)
plt.figure(figsize=(7,3))
plt.imshow(wordcloud, interpolation='bilinear')
plt.axis('off')
plt.title('Customer Feedback Wordcloud')
plt.show()
No description has been provided for this image
# Advanced: Cluster customers by tenure and monthly charges
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
clust_df = df[['tenure','MonthlyCharges']].dropna()
scaler = StandardScaler()
scaled = scaler.fit_transform(clust_df)
kmeans = KMeans(n_clusters=3, random_state=42)
labels = kmeans.fit_predict(scaled)
clust_df['Cluster'] = labels
plt.figure(figsize=(7,4))
sns.scatterplot(x='tenure', y='MonthlyCharges', hue='Cluster', data=clust_df, palette='viridis')
plt.title('Market Segmentation by Tenure & Charges')
plt.xlabel('Tenure (Months)')
plt.ylabel('Monthly Charges ($)')
plt.legend(title='Cluster')
plt.tight_layout()
plt.show()
No description has been provided for this image
# Generate synthetic cohort dataset
np.random.seed(0)
dates = pd.date_range('2021-01-01', periods=24, freq='ME')
cohort_df = pd.DataFrame({
    'CustomerID': np.random.randint(1000,2000,len(dates)),
    'Signup_Month': dates,
    'Active_Users': np.random.randint(50,300,len(dates))
})
print(cohort_df.head(3))
   CustomerID Signup_Month  Active_Users
0        1684   2021-01-31           138
1        1559   2021-02-28           131
2        1629   2021-03-31           215
# Line plot: Cohort Retention Trend
plt.figure(figsize=(8,4))
plt.plot(cohort_df['Signup_Month'], cohort_df['Active_Users'], marker='s', color='#b2182b')
plt.title('Monthly Active Users by Cohort Signup Month')
plt.xlabel('Signup Month')
plt.ylabel('Active Users')
plt.grid(True, ls=':', alpha=0.4)
plt.tight_layout()
plt.show()
No description has been provided for this image
# Detect missing values in survey data
missing = df.isnull().sum()
print('Missing survey responses per column:')
print(missing[missing > 0])
Missing survey responses per column:
Series([], dtype: int64)
# Visualize missing survey data
import missingno as msno
msno.matrix(df.sample(500, random_state=42))
plt.title('Missing Survey Data Visualization')
plt.show()
No description has been provided for this image
# Example: Incorrect aggregation warning
try:
    avg = df.groupby('Churn')['MonthlyCharges'].sum()['Yes'] / df['Churn'].value_counts()['Yes']
    print('Average monthly charge (churned):', avg)
except Exception as e:
    print('Aggregation error:', e)
Average monthly charge (churned): 74.44133226324237
# Example: NPS scale misinterpretation
promoters = (nps_df['NPS_Score'] >= 9).sum()
detractors = (nps_df['NPS_Score'] <= 6).sum()
nps_score = 100 * (promoters - detractors) / len(nps_df)
print('Net Promoter Score (NPS):', round(nps_score, 1))
Net Promoter Score (NPS): -47.6
# Best Practice: Cross-tabulation of Churn by Contract Type
ctab = pd.crosstab(df['Contract'], df['Churn'])
print(ctab)
Churn             No   Yes
Contract                  
Month-to-month  2220  1655
One year        1307   166
Two year        1647    48
# Build a satisfaction score index from available indicators
df['Satisfaction_Score'] = (df['tenure']/df['tenure'].max())*0.4 + (df['MonthlyCharges']/df['MonthlyCharges'].max())*0.6
print(df[['tenure','MonthlyCharges','Satisfaction_Score']].head(3))
   tenure  MonthlyCharges  Satisfaction_Score
0       1           29.85            0.156377
1      34           56.95            0.476636
2       2           53.85            0.283195
# Setup: Load bank marketing campaign dataset
dataset2 = openml.datasets.get_dataset(1461)
df2, _, _, _ = dataset2.get_data(dataset_format='dataframe')
df2.columns = ['age','job','marital','education','default','balance','housing','loan','contact','day','month','duration','campaign','pdays','previous','poutcome','response']
month_summary = df2['month'].value_counts().sort_index()
plt.figure(figsize=(7,3))
sns.lineplot(x=month_summary.index, y=month_summary.values, marker='o', color='#5e4fa2')
plt.title('Campaign Calls by Month')
plt.xlabel('Month')
plt.ylabel('Number of Calls')
plt.grid(ls=':', alpha=0.4)
plt.tight_layout()
plt.show()
No description has been provided for this image
# Mini Challenge: Identify if senior citizens are more likely to churn
senior_table = pd.crosstab(df['SeniorCitizen'], df['Churn'])
print('Senior Citizen by Churn Table:')
print(senior_table)
senior_churn_rate = senior_table.loc[1,'Yes'] / senior_table.loc[1].sum()
non_senior_churn_rate = senior_table.loc[0,'Yes'] / senior_table.loc[0].sum()
print('Churn Rate (Senior):', round(senior_churn_rate*100,1),'%')
print('Churn Rate (Non-Senior):', round(non_senior_churn_rate*100,1),'%')
Senior Citizen by Churn Table:
Churn            No   Yes
SeniorCitizen            
0              4508  1393
1               666   476
Churn Rate (Senior): 41.7 %
Churn Rate (Non-Senior): 23.6 %
# Save current figure as a PDF for business reporting
fig, ax = plt.subplots(figsize=(6,3))
sns.countplot(x='Contract', data=df, palette='Set2', ax=ax)
ax.set_title('Distribution of Contract Types')
plt.tight_layout()
fig.savefig('contract_types_distribution.pdf')
No description has been provided for this image
 

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.