Mathew K Analytics

Lesson 54 · Mastering Pandas

How to Encode Categorical Variables in Pandas for Effective Data Preparation

In this lesson, we will learn how to work with categorical data in pandas and transform text labels into numerical codes using practical, real-world…

⬇ Download notebookOpen in Colab ↗

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Encoding Categorical Variables in Pandas#

In this lesson, we will learn how to work with categorical data in pandas and transform text labels into numerical codes using practical, real-world datasets.

By the end, you will be able to choose the right method for encoding, know when to use each, and avoid common pitfalls.

# Suppress warnings for clean outputs
import warnings
import numpy as np
np.random.seed(42)
warnings.filterwarnings('ignore')
# Data setup (Titanic Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
(891, 12)
   PassengerId  Survived  Pclass  \
0            1         0       3   
1            2         1       1   
2            3         1       3   

                                                Name     Sex   Age  SibSp  \
0                            Braund, Mr. Owen Harris    male  22.0      1   
1  Cumings, Mrs. John Bradley (Florence Briggs Th...  female  38.0      1   
2                             Heikkinen, Miss. Laina  female  26.0      0   

   Parch            Ticket     Fare Cabin Embarked  
0      0         A/5 21171   7.2500   NaN        S  
1      0          PC 17599  71.2833   C85        C  
2      0  STON/O2. 3101282   7.9250   NaN        S  

What Are Categorical Variables?#

Categorical variables are columns with a limited number of values, usually text labels.

For example, 'Sex' and 'Embarked' in Titanic are categorical because they use labels like 'male', 'female', 'S', 'C', or 'Q' instead of numbers.

# List object-type columns in Titanic
cat_cols = df.select_dtypes(include=['object']).columns
print(cat_cols.tolist())
['Name', 'Sex', 'Ticket', 'Cabin', 'Embarked']
# Count the unique values in each categorical column
for col in cat_cols:
    print(f'{col}: {df[col].nunique()}')
    
Name: 891
Sex: 2
Ticket: 681
Cabin: 147
Embarked: 3

Why Encode Categorical Data?#

Most machine learning models require numbers, not text. Encoding categorical variables turns words into digits, letting algorithms learn from them.

There are two main strategies: label encoding and one-hot encoding.

# Label encoding simple categories: 'Sex'
df['Sex_encoded'] = df['Sex'].astype('category').cat.codes
print(df[['Sex', 'Sex_encoded']].head(6))
      Sex  Sex_encoded
0    male            1
1  female            0
2  female            0
3  female            0
4    male            1
5    male            1
# Map categories to custom values with a dictionary for 'Embarked'
embarked_map = {'S': 0, 'C': 1, 'Q': 2}
df['Embarked_encoded'] = df['Embarked'].map(embarked_map)
print(df[['Embarked', 'Embarked_encoded']].head(6))
  Embarked  Embarked_encoded
0        S               0.0
1        C               1.0
2        S               0.0
3        S               0.0
4        S               0.0
5        Q               2.0

One-Hot Encoding: When to Use It?#

One-hot encoding creates a separate column for each category, using 1 for present and 0 for absent.

It avoids implying order or priority, which is important for unordered categories.

# One-hot encode the 'Embarked' column
embarked_dummies = pd.get_dummies(df['Embarked'], prefix='Embarked')
print(embarked_dummies.head(6))
   Embarked_C  Embarked_Q  Embarked_S
0       False       False        True
1        True       False       False
2       False       False        True
3       False       False        True
4       False       False        True
5       False        True       False
# Join one-hot encoded columns back to our DataFrame
df = pd.concat([df, embarked_dummies], axis=1)
print(df.head(3))
   PassengerId  Survived  Pclass  \
0            1         0       3   
1            2         1       1   
2            3         1       3   

                                                Name     Sex   Age  SibSp  \
0                            Braund, Mr. Owen Harris    male  22.0      1   
1  Cumings, Mrs. John Bradley (Florence Briggs Th...  female  38.0      1   
2                             Heikkinen, Miss. Laina  female  26.0      0   

   Parch            Ticket     Fare Cabin Embarked  Sex_encoded  \
0      0         A/5 21171   7.2500   NaN        S            1   
1      0          PC 17599  71.2833   C85        C            0   
2      0  STON/O2. 3101282   7.9250   NaN        S            0   

   Embarked_encoded  Embarked_C  Embarked_Q  Embarked_S  
0               0.0       False       False        True  
1               1.0        True       False       False  
2               0.0       False       False        True  
# Encoding multiple categorical variables at once
cols_to_encode = ['Sex', 'Embarked']
df_multi = pd.get_dummies(df, columns=cols_to_encode)
print(df_multi.head(3))
   PassengerId  Survived  Pclass  \
0            1         0       3   
1            2         1       1   
2            3         1       3   

                                                Name   Age  SibSp  Parch  \
0                            Braund, Mr. Owen Harris  22.0      1      0   
1  Cumings, Mrs. John Bradley (Florence Briggs Th...  38.0      1      0   
2                             Heikkinen, Miss. Laina  26.0      0      0   

             Ticket     Fare Cabin  Sex_encoded  Embarked_encoded  Embarked_C  \
0         A/5 21171   7.2500   NaN            1               0.0       False   
1          PC 17599  71.2833   C85            0               1.0        True   
2  STON/O2. 3101282   7.9250   NaN            0               0.0       False   

   Embarked_Q  Embarked_S  Sex_female  Sex_male  Embarked_C  Embarked_Q  \
0       False        True       False      True       False       False   
1       False       False        True     False        True       False   
2       False        True        True     False       False       False   

   Embarked_S  
0        True  
1       False  
2        True  
# Handle missing data before encoding
df['Embarked_fill'] = df['Embarked'].fillna('Unknown')
print(df[['Embarked', 'Embarked_fill']].head(8))
  Embarked Embarked_fill
0        S             S
1        C             C
2        S             S
3        S             S
4        S             S
5        Q             Q
6        S             S
7        S             S
# One-hot encode including missing 'Unknown' level
embarked_full_dummies = pd.get_dummies(df['Embarked_fill'], prefix='Embarked')
print(embarked_full_dummies.head(8))
   Embarked_C  Embarked_Q  Embarked_S  Embarked_Unknown
0       False       False        True             False
1        True       False       False             False
2       False       False        True             False
3       False       False        True             False
4       False       False        True             False
5       False        True       False             False
6       False       False        True             False
7       False       False        True             False

Categorical Data Type for Efficiency#

Pandas offers a special 'category' dtype which saves space and speeds up some operations. It is very useful for columns with a handful of distinct values.

# Convert 'Embarked' to pandas 'category' type
df['Embarked_category'] = df['Embarked'].astype('category')
print(df['Embarked_category'].dtype)
print(df['Embarked_category'].memory_usage(deep=True))
category
1281
# Detect and encode ordinal categories
deck_map = {'A':1, 'B':2, 'C':3, 'D':4, 'E':5, 'F':6, 'G':7, 'Unknown':0}
df['Deck'] = df['Cabin'].str[0].fillna('Unknown')
df['Deck_encoded'] = df['Deck'].map(deck_map)
print(df[['Deck', 'Deck_encoded']].drop_duplicates().sort_values('Deck_encoded'))
        Deck  Deck_encoded
0    Unknown           0.0
23         A           1.0
31         B           2.0
1          C           3.0
21         D           4.0
6          E           5.0
66         F           6.0
10         G           7.0
339        T           NaN
# Practice: Can you encode the 'Pclass' column as category?
df['Pclass_cat'] = df['Pclass'].astype('category')
print(df[['Pclass', 'Pclass_cat']].head(6))
   Pclass Pclass_cat
0       3          3
1       1          1
2       3          3
3       1          1
4       3          3
5       3          3

Troubleshooting Encoding Problems#

  • Nulls or spelling issues can break encoding.
  • High-cardinality columns like 'Name' are not good for one-hot encoding.
  • Recheck for unexpected new categories before predicting!

Recap#

  • Label encoding is best for clear, ordered categories.
  • One-hot encoding is right for unordered choices.
  • Always handle missing data before encoding.
  • The category dtype economizes memory.

Ready for a real challenge? Try encoding 'Ticket' using frequency or target-based methods!

Keep Learning and Connect!#

If this helped you out, remember to like and subscribe for more practical pandas walkthroughs. Leave a comment sharing how you use encoding in your own projects!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.