Lesson 15 · Market Research Analytics in Python
Assessing Data Quality in Market Research
In this lesson, we will solve a real-world problem: How do we assess and improve data quality in customer surveys and market research datasets? Many…
- CourseMarket Research Analytics in Python
- Lesson15 of 56
- Video23 min
- FormatJupyter notebook · 24 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbAssessing Data Quality in Market Research#
- In this lesson, we will solve a real-world problem: How do we assess and improve data quality in customer surveys and market research datasets?
- Many business decisions rely on survey or customer data; poor data quality can lead to wrong conclusions.
- You will learn how to explore, diagnose, and handle common data quality issues using Python.
- By the end, you will produce actionable insights by ensuring your data is reliable before analysis.
import warnings
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import openml
warnings.filterwarnings('ignore')
What is market research data, and why does quality matter?#
- Market research data comes from customer surveys, feedback, sales, or marketing experiments.
- Most datasets include multiple columns: respondent demographics, product ratings, open-ended feedback, or behavioral signals.
- Beginners often forget to check for mistakes like missing data, inconsistent codes, or survey fatigue, which can lead to misleading analysis.
- Good data quality is the foundation for accurate, actionable business insights.
# Beginner: Load a real customer satisfaction survey dataset
dataset = openml.datasets.get_dataset(42178)
df = dataset.get_data(dataset_format='dataframe')[0]
print(df.shape)
print(df.head(3))
Beginner: Which columns and customer features are in the data?#
- The customer satisfaction survey includes demographic and service columns such as gender, SeniorCitizen, tenure, PaymentMethod, and Churn.
- Each row represents a unique customer record; columns represent survey responses or customer attributes.
- Identifying the variables in your data is the first step to understanding data quality.
# Beginner: Check for missing values in the survey data
missing_counts = df.isnull().sum()
print(missing_counts[missing_counts > 0])
# Beginner: Count unique values in Churn column
print('Unique churn responses:', df['Churn'].unique())
# Check and visualize missing values safely
missing = df.isnull().sum().sort_values(ascending=False)
missing_nonzero = missing[missing > 0]
if len(missing_nonzero) > 0:
plt.figure(figsize=(10,4))
missing_nonzero.plot(kind='bar', color='salmon')
plt.title('Missing Value Count by Column')
plt.xlabel('Column')
plt.ylabel('Count')
plt.show()
else:
print('No missing values found in the dataset.')
# Beginner: Load a marketing campaign dataset
dataset = openml.datasets.get_dataset(1461)
df_marketing, _, _, _ = dataset.get_data(dataset_format='dataframe')
df_marketing.columns = ['age','job','marital','education','default','balance','housing','loan','contact','day','month','duration','campaign','pdays','previous','poutcome','response']
print(df_marketing.shape)
print(df_marketing.head(3))
# Intermediate: Look for duplicate records in the marketing dataset
duplicates = df_marketing.duplicated().sum()
print(f'Duplicate records found: {duplicates}')
# Intermediate: Assess data type consistency in each column
print(df_marketing.dtypes)
# Intermediate: Frequency of each unique value in the 'response' column
response_counts = df_marketing['response'].value_counts(dropna=False)
print(response_counts)
# Intermediate: Spot potential outliers in 'age' and 'balance' fields
print('Age range:', df_marketing['age'].min(), '-', df_marketing['age'].max())
print('Balance range:', df_marketing['balance'].min(), '-', df_marketing['balance'].max())
# Intermediate: Visualize 'age' distribution to spot unusual spikes or cuts
plt.figure(figsize=(8,4))
df_marketing['age'].hist(bins=30, color='skyblue', edgecolor='black')
plt.title('Customer Age Distribution')
plt.xlabel('Age')
plt.ylabel('Count')
plt.show()
# Advanced: Explore data quality in open feedback text
df_feedback = pd.DataFrame({'CustomerID':[1,2,3,4,5], 'Feedback':['Great service and friendly staff','Delivery was slow and packaging was poor','Excellent quality, will buy again','Customer support needs improvement','Good value for money']})
print(df_feedback.head(3))
print('Feedback length:', df_feedback['Feedback'].apply(len))
# Advanced: Check for non-English or gibberish feedback (simple heuristic)
def is_english(text):
try:
text.encode(encoding='utf-8').decode('ascii')
except UnicodeDecodeError:
return False
return True
non_english = df_feedback[~df_feedback['Feedback'].apply(is_english)]
print('Non-English or gibberish feedback:')
print(non_english)
# Advanced: Load Net Promoter Score (NPS) survey and check for invalid scores
np.random.seed(42)
df_nps = pd.DataFrame({'CustomerID': range(1,501), 'Age': np.random.randint(18,70,500), 'Region': np.random.choice(['North','South','East','West'],500), 'NPS_Score': np.random.randint(0,11,500)})
print(df_nps.head(3))
invalid_nps = df_nps[~df_nps['NPS_Score'].between(0,10)]
print('Invalid NPS records:', invalid_nps.shape[0])
# Advanced: Visualize NPS score distribution to spot survey response artifacts
plt.figure(figsize=(7,3))
df_nps['NPS_Score'].hist(bins=11, color='green', rwidth=0.8)
plt.xticks(range(0,11))
plt.title('NPS Survey Score Distribution')
plt.xlabel('NPS Score')
plt.ylabel('Number of Responses')
plt.show()
# Error Handling: Simulate missing survey responses in key columns
df_missing = df.copy()
df_missing.loc[df_missing.sample(frac=0.1, random_state=42).index, 'MonthlyCharges'] = np.nan
print('Missing MonthlyCharges:', df_missing['MonthlyCharges'].isnull().sum())
# Error Handling: Attempt incorrect grouping on a continuous variable
try:
df.groupby('MonthlyCharges').size()
except Exception as e:
print('Grouping error:', e)
# Error Handling: Misinterpret Likert or NPS scales (e.g., average as a metric)
mean_nps = df_nps['NPS_Score'].mean()
print('Average NPS (should not be used as standard NPS):', mean_nps)
Best Practices: Segmentation, cross-tabs, index construction, trend analysis#
- Segmenting by customer age, tenure, or region uncovers actionable subgroups.
- Cross-tabulation helps reveal relationships between two categorical fields, such as churn and contract type.
- Building custom indexes (such as satisfaction scores) and running trends over time is critical for decision support.
- Always start with data quality before segmentation, as bad data can hide or create fake patterns!
# Pattern: Segment NPS by Region
nps_by_region = df_nps.groupby('Region')['NPS_Score'].mean()
print(nps_by_region)
# Pattern: Build a cross-tab between Churn and Contract type
churn_contract = pd.crosstab(df['Churn'], df['Contract'])
print(churn_contract)
# Pattern: Build a custom satisfaction index from multiple fields
for col in ['OnlineSecurity', 'OnlineBackup', 'TechSupport']:
df[col] = df[col].map({'Yes':1, 'No':0})
df['Satisfaction_Index'] = df[['OnlineSecurity','OnlineBackup','TechSupport']].mean(axis=1)
print(df['Satisfaction_Index'].head(5))
# Pattern: Investigate churn trends over customer tenure
churned_by_tenure = df.groupby('tenure')['Churn'].value_counts(normalize=True).unstack().fillna(0)
churned_by_tenure.plot(kind='line', figsize=(10,4))
plt.ylabel('Proportion of Churn')
plt.title('Churn Proportion by Customer Tenure')
plt.show()
# End-to-end: From raw survey data to usable business insight
# Step 1: Load customer satisfaction data
df_end = df.copy()
# Step 2: Clean Churn column (remove whitespaces, fix typos if any)
df_end['Churn'] = df_end['Churn'].str.strip().str.title()
# Step 3: Impute missing MonthlyCharges using median
df_end['MonthlyCharges'] = df_end['MonthlyCharges'].fillna(df_end['MonthlyCharges'].median())
# Step 4: Compute the churn rate
churn_rate = (df_end['Churn'] == 'Yes').mean()
print(f'Cleaned dataset churn rate: {churn_rate:.2%}')
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



