Mathew K Analytics

Lesson 52 · Mastering Pandas

Master Feature Engineering with Pandas for Powerful Data Preprocessing

Welcome! Today we will explore how to use pandas for feature engineering, a core skill in data analysis and machine learning. Feature engineering is about…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Feature Engineering with Pandas#

Welcome! Today we will explore how to use pandas for feature engineering, a core skill in data analysis and machine learning.

Feature engineering is about creating new data columns (features) that help models and humans understand the data better.

We will use the Titanic dataset to practice each new concept step by step.

import warnings; warnings.filterwarnings("ignore")
import pandas as pd
import numpy as np
np.random.seed(42)
# Data setup (Titanic Dataset)
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
(891, 12)
   PassengerId  Survived  Pclass  \
0            1         0       3   
1            2         1       1   
2            3         1       3   

                                                Name     Sex   Age  SibSp  \
0                            Braund, Mr. Owen Harris    male  22.0      1   
1  Cumings, Mrs. John Bradley (Florence Briggs Th...  female  38.0      1   
2                             Heikkinen, Miss. Laina  female  26.0      0   

   Parch            Ticket     Fare Cabin Embarked  
0      0         A/5 21171   7.2500   NaN        S  
1      0          PC 17599  71.2833   C85        C  
2      0  STON/O2. 3101282   7.9250   NaN        S  

Why Engineer Features?#

Raw data almost never has all the information we need.

By combining, changing, or extracting values, we help computers and ourselves make better predictions.

On Titanic, for example, we can guess who had better chances of survival by making new features from names or families.

# Creating a new column: FamilySize
df['FamilySize'] = df['SibSp'] + df['Parch'] + 1
print(df[['SibSp', 'Parch', 'FamilySize']].head())
   SibSp  Parch  FamilySize
0      1      0           2
1      1      0           2
2      0      0           1
3      1      0           2
4      0      0           1
# Extract title from Name column to create a new feature
df['Title'] = df['Name'].str.extract(' ([A-Za-z]+)\.')
print(df[['Name', 'Title']].head())
                                                Name Title
0                            Braund, Mr. Owen Harris    Mr
1  Cumings, Mrs. John Bradley (Florence Briggs Th...   Mrs
2                             Heikkinen, Miss. Laina  Miss
3       Futrelle, Mrs. Jacques Heath (Lily May Peel)   Mrs
4                           Allen, Mr. William Henry    Mr
# Create AgeGroup column: Bin ages into categories
bins = [0, 12, 18, 35, 60, 99]
labels = ['Child', 'Teen', 'Adult', 'Middle Aged', 'Senior']
df['AgeGroup'] = pd.cut(df['Age'], bins=bins, labels=labels)
print(df[['Age', 'AgeGroup']].head(10))
    Age     AgeGroup
0  22.0        Adult
1  38.0  Middle Aged
2  26.0        Adult
3  35.0        Adult
4  35.0        Adult
5   NaN          NaN
6  54.0  Middle Aged
7   2.0        Child
8  27.0        Adult
9  14.0         Teen
# Handle missing ages by filling with AgeGroup median
medians = df.groupby('AgeGroup')['Age'].transform('median')
df['AgeFilled'] = df['Age'].fillna(medians)
print(df.loc[df['Age'].isnull(), ['AgeGroup', 'Age', 'AgeFilled']].head())
   AgeGroup  Age  AgeFilled
5       NaN  NaN        NaN
17      NaN  NaN        NaN
19      NaN  NaN        NaN
26      NaN  NaN        NaN
28      NaN  NaN        NaN
# Create IsAlone feature: 1 if passenger is alone, else 0
df['IsAlone'] = (df['FamilySize'] == 1).astype(int)
print(df[['FamilySize', 'IsAlone']].head(8))
   FamilySize  IsAlone
0           2        0
1           2        0
2           1        1
3           2        0
4           1        1
5           1        1
6           1        1
7           5        0
# Make a categorical variable: Deck extracted from Cabin
df['Deck'] = df['Cabin'].str[0]
print(df[['Cabin', 'Deck']].head(8))
  Cabin Deck
0   NaN  NaN
1   C85    C
2   NaN  NaN
3  C123    C
4   NaN  NaN
5   NaN  NaN
6   E46    E
7   NaN  NaN
# Create Fare per person: Fare divided by FamilySize
df['FarePerPerson'] = df['Fare'] / df['FamilySize']
print(df[['Fare', 'FamilySize', 'FarePerPerson']].head(8))
      Fare  FamilySize  FarePerPerson
0   7.2500           2        3.62500
1  71.2833           2       35.64165
2   7.9250           1        7.92500
3  53.1000           2       26.55000
4   8.0500           1        8.05000
5   8.4583           1        8.45830
6  51.8625           1       51.86250
7  21.0750           5        4.21500
# Map Sex column to numeric for analysis
df['SexNum'] = df['Sex'].map({'male': 0, 'female': 1})
print(df[['Sex', 'SexNum']].head(8))
      Sex  SexNum
0    male       0
1  female       1
2  female       1
3  female       1
4    male       0
5    male       0
6    male       0
7    male       0
# One-hot encode Embarked column
embark_dummies = pd.get_dummies(df['Embarked'], prefix='Embarked')
df = pd.concat([df, embark_dummies], axis=1)
print(df[['Embarked', 'Embarked_C', 'Embarked_Q', 'Embarked_S']].head(8))
  Embarked  Embarked_C  Embarked_Q  Embarked_S
0        S       False       False        True
1        C        True       False       False
2        S       False       False        True
3        S       False       False        True
4        S       False       False        True
5        Q       False        True       False
6        S       False       False        True
7        S       False       False        True
# Create interaction feature: Pclass by Sex
df['ClassSex'] = df['Pclass'].astype(str) + '_' + df['Sex']
print(df[['Pclass', 'Sex', 'ClassSex']].head(8))
   Pclass     Sex  ClassSex
0       3    male    3_male
1       1  female  1_female
2       3  female  3_female
3       1  female  1_female
4       3    male    3_male
5       3    male    3_male
6       1    male    1_male
7       3    male    3_male
# Detect outliers in Fare using a Z-score method
fare_mean = df['Fare'].mean()
fare_std = df['Fare'].std()
df['FareZscore'] = (df['Fare'] - fare_mean) / fare_std
print(df[['Fare', 'FareZscore']].head(8))
      Fare  FareZscore
0   7.2500   -0.502163
1  71.2833    0.786404
2   7.9250   -0.488580
3  53.1000    0.420494
4   8.0500   -0.486064
5   8.4583   -0.477848
6  51.8625    0.395591
7  21.0750   -0.223957

Mini-Project: Feature Engineering EDA#

Let us pause and practice turning our new features into insight.

We will ask if traveling alone or with family changed survival rates.

Try to guess before running the code below!

# Group by IsAlone and calculate mean survival
print(df.groupby('IsAlone')['Survived'].mean())
IsAlone
0    0.505650
1    0.303538
Name: Survived, dtype: float64
# Average survival by AgeGroup
print(df.groupby('AgeGroup')['Survived'].mean())
AgeGroup
Child          0.579710
Teen           0.428571
Adult          0.382682
Middle Aged    0.400000
Senior         0.227273
Name: Survived, dtype: float64
# Visualization: Survival by AgeGroup and Sex
import matplotlib.pyplot as plt
pd.crosstab(df['AgeGroup'], df['Sex']).plot(kind='bar', stacked=True, figsize=(6,4))
plt.title('Count of Passengers by Age Group and Sex')
plt.xlabel('Age Group')
plt.ylabel('Count')
plt.legend(title='Sex')
plt.tight_layout()
plt.show()
No description has been provided for this image
# Quick correlation check for new features
new_features = ['FamilySize', 'IsAlone', 'SexNum', 'FarePerPerson']
print(df[new_features + ['Survived']].corr())
               FamilySize   IsAlone    SexNum  FarePerPerson  Survived
FamilySize       1.000000 -0.690922  0.200988      -0.099173  0.016639
IsAlone         -0.690922  1.000000 -0.303646       0.045603 -0.203367
SexNum           0.200988 -0.303646  1.000000       0.115143  0.543351
FarePerPerson   -0.099173  0.045603  0.115143       1.000000  0.221600
Survived         0.016639 -0.203367  0.543351       0.221600  1.000000
# Caution: Avoid data leakage
df['SurvivalKnown'] = df['Survived'] # This would leak the target.
print('Columns:', df.columns.tolist())
del df['SurvivalKnown']
Columns: ['PassengerId', 'Survived', 'Pclass', 'Name', 'Sex', 'Age', 'SibSp', 'Parch', 'Ticket', 'Fare', 'Cabin', 'Embarked', 'FamilySize', 'Title', 'AgeGroup', 'AgeFilled', 'IsAlone', 'Deck', 'FarePerPerson', 'SexNum', 'Embarked_C', 'Embarked_Q', 'Embarked_S', 'ClassSex', 'FareZscore', 'SurvivalKnown']

Recap: What We Learned#

  • Creating features by combining columns
  • Extracting information from text
  • Handling missing values
  • Making categories
  • Encoding labels for models
  • Checking for outliers and leakage

Practice these steps to improve any data project!

Keep Practicing & Subscribe!#

Want more pandas tutorials and tips?

Practice: Pick a fun dataset, design and test at least two of your own features.

Subscribe for future walkthroughs on pandas, data science, and machine learning!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.