Lesson 19 · Market Research Analytics in Python
Encoding Categorical and Survey Variables in Python | Market Research Analytics Tutorial
Learn why survey and customer data often require encoding for analysis Understand why proper encoding drives valid business insights Explore how encoded…
- CourseMarket Research Analytics in Python
- Lesson19 of 56
- Video18 min
- FormatJupyter notebook · 20 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbEncoding Categorical and Survey Variables for Market Research Analytics#
- Learn why survey and customer data often require encoding for analysis
- Understand why proper encoding drives valid business insights
- Explore how encoded variables improve segmentation and modeling
- Practice with real datasets: customer satisfaction, NPS, and more
- By the end, you will be able to turn raw survey responses into actionable metrics
import pandas as pd
import numpy as np
import openml
import warnings
warnings.filterwarnings('ignore')
Core concepts: Categorical and Survey Variables in Customer Analytics#
- Market research data includes: demographics, survey responses, feedback, and more
- Categorical variables capture types (e.g. payment method, gender, region)
- Survey variables can be ordinal (ratings, Likert scales), nominal (choices), or free text
- Beginners often misinterpret text, forget to encode categories, or treat ratings as equally spaced
# Load a real customer satisfaction dataset from OpenML
dataset = openml.datasets.get_dataset(42178)
df = dataset.get_data(dataset_format='dataframe')[0]
print(df.shape)
print(df.head(3))
# Check the data types of each column
print(df.dtypes)
# Encode a binary categorical variable: 'gender'
df['gender_encoded'] = df['gender'].map({'Male': 1, 'Female': 0})
print(df[['gender', 'gender_encoded']].head(5))
# One-hot encode a nominal variable: 'PaymentMethod'
payment_dummies = pd.get_dummies(df['PaymentMethod'], prefix='PayMethod')
df = pd.concat([df, payment_dummies], axis=1)
print(payment_dummies.head(3))
# Encode an ordinal variable: 'Contract' has an order of importance
contract_order = {'Month-to-month': 0, 'One year': 1, 'Two year': 2}
df['contract_encoded'] = df['Contract'].map(contract_order)
print(df[['Contract', 'contract_encoded']].head(5))
# Example: Encoding survey answers on a Likert scale (Strongly disagree to Strongly agree)
survey_map = {'Strongly disagree': 1, 'Disagree': 2, 'Neutral': 3, 'Agree': 4, 'Strongly agree': 5}
example_survey = pd.Series(['Strongly agree', 'Neutral', 'Agree', 'Disagree', 'Agree'])
encoded_survey = example_survey.map(survey_map)
print(pd.DataFrame({'Original': example_survey, 'Encoded': encoded_survey}))
# Load a synthetic NPS (Net Promoter Score) dataset for practice
np.random.seed(42)
nps_df = pd.DataFrame({
'CustomerID': range(1, 501),
'Age': np.random.randint(18, 70, 500),
'Region': np.random.choice(['North', 'South', 'East', 'West'], 500),
'NPS_Score': np.random.randint(0, 11, 500)
})
print(nps_df.head(3))
# Turn NPS scores into Promoter, Passive, Detractor categories
def nps_category(score):
if score <= 6:
return 'Detractor'
elif score <= 8:
return 'Passive'
else:
return 'Promoter'
nps_df['NPS_Type'] = nps_df['NPS_Score'].apply(nps_category)
print(nps_df[['NPS_Score', 'NPS_Type']].head(7))
# Ordinal encode 'NPS_Type' to help with correlation analysis
nps_order = {'Detractor': 0, 'Passive': 1, 'Promoter': 2}
nps_df['NPS_Type_Encoded'] = nps_df['NPS_Type'].map(nps_order)
print(nps_df[['NPS_Type', 'NPS_Type_Encoded']].head(7))
# One-hot encode 'Region' variable in the NPS dataset
region_dummies = pd.get_dummies(nps_df['Region'], prefix='Region')
nps_df = pd.concat([nps_df, region_dummies], axis=1)
print(region_dummies.head(5))
# ERROR HANDLING: Find missing survey responses in the main dataset
missing = df.isnull().sum()
print('Missing values per column:')
print(missing[missing > 0])
# ERROR HANDLING: Try to encode a column with unexpected string values
try:
df['SeniorCitizenNumeric'] = df['SeniorCitizen'].astype(int)
except Exception as e:
print('Error:', e)
print('Try using a map or replace on unique values before encoding.')
print('Unique values:', df['SeniorCitizen'].unique())
# ERROR HANDLING: Mixing up Likert scales ordering
wrong_map = {'Strongly disagree': 5, 'Disagree': 4, 'Neutral': 3, 'Agree': 2, 'Strongly agree': 1}
example_wrong = pd.Series(['Agree', 'Strongly agree', 'Disagree'])
encoded_wrong = example_wrong.map(wrong_map)
print(pd.DataFrame({'Original': example_wrong, 'Incorrect_Encoding': encoded_wrong}))
print('Warning: This reverses the real meaning of the answers! Always double-check order.')
# Best practice: Check value counts before and after encoding
print('Before encoding - PaymentMethod counts:')
print(df['PaymentMethod'].value_counts())
print('One-hot encoded columns summary:')
print(df.filter(like='PayMethod').sum())
# Cross-tabulate encoded survey responses by region (NPS example)
region_nps = pd.crosstab(nps_df['Region'], nps_df['NPS_Type'])
print(region_nps)
# ADVANCED: Segment customers using multiple encoded columns
segments = df.groupby(['gender_encoded', 'contract_encoded'])['MonthlyCharges'].mean().unstack()
print(segments)
# ADVANCED: Build a simple index score from encoded variables
df['ServiceScore'] = df[['OnlineSecurity', 'OnlineBackup', 'TechSupport']].apply(lambda x: sum(x == 'Yes'), axis=1)
print(df[['OnlineSecurity', 'OnlineBackup', 'TechSupport', 'ServiceScore']].head(5))
Best practices for survey variable encoding#
- Always review unique values before choosing encoding methods
- Match business meaning: ordinal vs. nominal vs. binary
- Validate encoding logic with value_counts or summary tables
- Document decisions for future analysis or collaboration
- Use encoded features for segmentation, trends, cross-tabs, and scoring
# END-TO-END: From raw survey to actionable customer insight
nps_summary = nps_df.groupby('NPS_Type').size().reset_index(name='Count')
nps_summary['Percent'] = 100 * nps_summary['Count'] / nps_summary['Count'].sum()
print(nps_summary)
if (nps_summary[nps_summary['NPS_Type']=='Detractor']['Percent'].values[0] > 20):
print('Recommendation: Take urgent action with Detractors before running another campaign.')
else:
print('Recommendation: Detractor share is under control; focus on converting Passives to Promoters.')
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



