Mathew K Analytics

Lesson 19 · Market Research Analytics in Python

Encoding Categorical and Survey Variables in Python | Market Research Analytics Tutorial

Learn why survey and customer data often require encoding for analysis Understand why proper encoding drives valid business insights Explore how encoded…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Encoding Categorical and Survey Variables for Market Research Analytics#

  • Learn why survey and customer data often require encoding for analysis
  • Understand why proper encoding drives valid business insights
  • Explore how encoded variables improve segmentation and modeling
  • Practice with real datasets: customer satisfaction, NPS, and more
  • By the end, you will be able to turn raw survey responses into actionable metrics
import pandas as pd
import numpy as np
import openml
import warnings
warnings.filterwarnings('ignore')

Core concepts: Categorical and Survey Variables in Customer Analytics#

  • Market research data includes: demographics, survey responses, feedback, and more
  • Categorical variables capture types (e.g. payment method, gender, region)
  • Survey variables can be ordinal (ratings, Likert scales), nominal (choices), or free text
  • Beginners often misinterpret text, forget to encode categories, or treat ratings as equally spaced
# Load a real customer satisfaction dataset from OpenML
dataset = openml.datasets.get_dataset(42178)
df = dataset.get_data(dataset_format='dataframe')[0]
print(df.shape)
print(df.head(3))
(7043, 20)
   gender  SeniorCitizen Partner Dependents  tenure PhoneService  \
0  Female              0     Yes         No       1           No   
1    Male              0      No         No      34          Yes   
2    Male              0      No         No       2          Yes   

      MultipleLines InternetService OnlineSecurity OnlineBackup  \
0  No phone service             DSL             No          Yes   
1                No             DSL            Yes           No   
2                No             DSL            Yes          Yes   

  DeviceProtection TechSupport StreamingTV StreamingMovies        Contract  \
0               No          No          No              No  Month-to-month   
1              Yes          No          No              No        One year   
2               No          No          No              No  Month-to-month   

  PaperlessBilling     PaymentMethod  MonthlyCharges TotalCharges Churn  
0              Yes  Electronic check           29.85        29.85    No  
1               No      Mailed check           56.95       1889.5    No  
2              Yes      Mailed check           53.85       108.15   Yes  
# Check the data types of each column
print(df.dtypes)
gender               object
SeniorCitizen         uint8
Partner              object
Dependents           object
tenure                uint8
PhoneService         object
MultipleLines        object
InternetService      object
OnlineSecurity       object
OnlineBackup         object
DeviceProtection     object
TechSupport          object
StreamingTV          object
StreamingMovies      object
Contract             object
PaperlessBilling     object
PaymentMethod        object
MonthlyCharges      float64
TotalCharges         object
Churn                object
dtype: object
# Encode a binary categorical variable: 'gender'
df['gender_encoded'] = df['gender'].map({'Male': 1, 'Female': 0})
print(df[['gender', 'gender_encoded']].head(5))
   gender  gender_encoded
0  Female               0
1    Male               1
2    Male               1
3    Male               1
4  Female               0
# One-hot encode a nominal variable: 'PaymentMethod'
payment_dummies = pd.get_dummies(df['PaymentMethod'], prefix='PayMethod')
df = pd.concat([df, payment_dummies], axis=1)
print(payment_dummies.head(3))
   PayMethod_Bank transfer (automatic)  PayMethod_Credit card (automatic)  \
0                                False                              False   
1                                False                              False   
2                                False                              False   

   PayMethod_Electronic check  PayMethod_Mailed check  
0                        True                   False  
1                       False                    True  
2                       False                    True  
# Encode an ordinal variable: 'Contract' has an order of importance
contract_order = {'Month-to-month': 0, 'One year': 1, 'Two year': 2}
df['contract_encoded'] = df['Contract'].map(contract_order)
print(df[['Contract', 'contract_encoded']].head(5))
         Contract  contract_encoded
0  Month-to-month                 0
1        One year                 1
2  Month-to-month                 0
3        One year                 1
4  Month-to-month                 0
# Example: Encoding survey answers on a Likert scale (Strongly disagree to Strongly agree)
survey_map = {'Strongly disagree': 1, 'Disagree': 2, 'Neutral': 3, 'Agree': 4, 'Strongly agree': 5}
example_survey = pd.Series(['Strongly agree', 'Neutral', 'Agree', 'Disagree', 'Agree'])
encoded_survey = example_survey.map(survey_map)
print(pd.DataFrame({'Original': example_survey, 'Encoded': encoded_survey}))
         Original  Encoded
0  Strongly agree        5
1         Neutral        3
2           Agree        4
3        Disagree        2
4           Agree        4
# Load a synthetic NPS (Net Promoter Score) dataset for practice
np.random.seed(42)
nps_df = pd.DataFrame({
    'CustomerID': range(1, 501),
    'Age': np.random.randint(18, 70, 500),
    'Region': np.random.choice(['North', 'South', 'East', 'West'], 500),
    'NPS_Score': np.random.randint(0, 11, 500)
})
print(nps_df.head(3))
   CustomerID  Age Region  NPS_Score
0           1   56   West          2
1           2   69  North          0
2           3   46   East          4
# Turn NPS scores into Promoter, Passive, Detractor categories
def nps_category(score):
    if score <= 6:
        return 'Detractor'
    elif score <= 8:
        return 'Passive'
    else:
        return 'Promoter'
nps_df['NPS_Type'] = nps_df['NPS_Score'].apply(nps_category)
print(nps_df[['NPS_Score', 'NPS_Type']].head(7))
   NPS_Score   NPS_Type
0          2  Detractor
1          0  Detractor
2          4  Detractor
3          3  Detractor
4          9   Promoter
5          7    Passive
6          0  Detractor
# Ordinal encode 'NPS_Type' to help with correlation analysis
nps_order = {'Detractor': 0, 'Passive': 1, 'Promoter': 2}
nps_df['NPS_Type_Encoded'] = nps_df['NPS_Type'].map(nps_order)
print(nps_df[['NPS_Type', 'NPS_Type_Encoded']].head(7))
    NPS_Type  NPS_Type_Encoded
0  Detractor                 0
1  Detractor                 0
2  Detractor                 0
3  Detractor                 0
4   Promoter                 2
5    Passive                 1
6  Detractor                 0
# One-hot encode 'Region' variable in the NPS dataset
region_dummies = pd.get_dummies(nps_df['Region'], prefix='Region')
nps_df = pd.concat([nps_df, region_dummies], axis=1)
print(region_dummies.head(5))
   Region_East  Region_North  Region_South  Region_West
0        False         False         False         True
1        False          True         False        False
2         True         False         False        False
3        False         False         False         True
4         True         False         False        False
# ERROR HANDLING: Find missing survey responses in the main dataset
missing = df.isnull().sum()
print('Missing values per column:')
print(missing[missing > 0])
Missing values per column:
Series([], dtype: int64)
# ERROR HANDLING: Try to encode a column with unexpected string values
try:
    df['SeniorCitizenNumeric'] = df['SeniorCitizen'].astype(int)
except Exception as e:
    print('Error:', e)
    print('Try using a map or replace on unique values before encoding.')
    print('Unique values:', df['SeniorCitizen'].unique())
# ERROR HANDLING: Mixing up Likert scales ordering
wrong_map = {'Strongly disagree': 5, 'Disagree': 4, 'Neutral': 3, 'Agree': 2, 'Strongly agree': 1}
example_wrong = pd.Series(['Agree', 'Strongly agree', 'Disagree'])
encoded_wrong = example_wrong.map(wrong_map)
print(pd.DataFrame({'Original': example_wrong, 'Incorrect_Encoding': encoded_wrong}))
print('Warning: This reverses the real meaning of the answers! Always double-check order.')
         Original  Incorrect_Encoding
0           Agree                   2
1  Strongly agree                   1
2        Disagree                   4
Warning: This reverses the real meaning of the answers! Always double-check order.
# Best practice: Check value counts before and after encoding
print('Before encoding - PaymentMethod counts:')
print(df['PaymentMethod'].value_counts())
print('One-hot encoded columns summary:')
print(df.filter(like='PayMethod').sum())
Before encoding - PaymentMethod counts:
PaymentMethod
Electronic check             2365
Mailed check                 1612
Bank transfer (automatic)    1544
Credit card (automatic)      1522
Name: count, dtype: int64
One-hot encoded columns summary:
PayMethod_Bank transfer (automatic)    1544
PayMethod_Credit card (automatic)      1522
PayMethod_Electronic check             2365
PayMethod_Mailed check                 1612
dtype: int64
# Cross-tabulate encoded survey responses by region (NPS example)
region_nps = pd.crosstab(nps_df['Region'], nps_df['NPS_Type'])
print(region_nps)
NPS_Type  Detractor  Passive  Promoter
Region                                
East             76       20        17
North            86       38        25
South            84       19        18
West             77       15        25
# ADVANCED: Segment customers using multiple encoded columns
segments = df.groupby(['gender_encoded', 'contract_encoded'])['MonthlyCharges'].mean().unstack()
print(segments)
contract_encoded          0          1          2
gender_encoded                                   
0                 66.652623  66.841643  60.513373
1                 66.147615  63.343444  61.025941
# ADVANCED: Build a simple index score from encoded variables
df['ServiceScore'] = df[['OnlineSecurity', 'OnlineBackup', 'TechSupport']].apply(lambda x: sum(x == 'Yes'), axis=1)
print(df[['OnlineSecurity', 'OnlineBackup', 'TechSupport', 'ServiceScore']].head(5))
  OnlineSecurity OnlineBackup TechSupport  ServiceScore
0             No          Yes          No             1
1            Yes           No          No             1
2            Yes          Yes          No             2
3            Yes           No         Yes             2
4             No           No          No             0

Best practices for survey variable encoding#

  • Always review unique values before choosing encoding methods
  • Match business meaning: ordinal vs. nominal vs. binary
  • Validate encoding logic with value_counts or summary tables
  • Document decisions for future analysis or collaboration
  • Use encoded features for segmentation, trends, cross-tabs, and scoring
# END-TO-END: From raw survey to actionable customer insight
nps_summary = nps_df.groupby('NPS_Type').size().reset_index(name='Count')
nps_summary['Percent'] = 100 * nps_summary['Count'] / nps_summary['Count'].sum()
print(nps_summary)
if (nps_summary[nps_summary['NPS_Type']=='Detractor']['Percent'].values[0] > 20):
    print('Recommendation: Take urgent action with Detractors before running another campaign.')
else:
    print('Recommendation: Detractor share is under control; focus on converting Passives to Promoters.')
    NPS_Type  Count  Percent
0  Detractor    323     64.6
1    Passive     92     18.4
2   Promoter     85     17.0
Recommendation: Take urgent action with Detractors before running another campaign.
 

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.