Lesson 52 · Mastering Pandas
Master Feature Engineering with Pandas for Powerful Data Preprocessing
Welcome! Today we will explore how to use pandas for feature engineering, a core skill in data analysis and machine learning. Feature engineering is about…
- CourseMastering Pandas
- Lesson52 of 44
- Video21 min
- FormatJupyter notebook · 17 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbFeature Engineering with Pandas#
Welcome! Today we will explore how to use pandas for feature engineering, a core skill in data analysis and machine learning.
Feature engineering is about creating new data columns (features) that help models and humans understand the data better.
We will use the Titanic dataset to practice each new concept step by step.
import warnings; warnings.filterwarnings("ignore")
import pandas as pd
import numpy as np
np.random.seed(42)
# Data setup (Titanic Dataset)
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
Why Engineer Features?#
Raw data almost never has all the information we need.
By combining, changing, or extracting values, we help computers and ourselves make better predictions.
On Titanic, for example, we can guess who had better chances of survival by making new features from names or families.
# Creating a new column: FamilySize
df['FamilySize'] = df['SibSp'] + df['Parch'] + 1
print(df[['SibSp', 'Parch', 'FamilySize']].head())
# Extract title from Name column to create a new feature
df['Title'] = df['Name'].str.extract(' ([A-Za-z]+)\.')
print(df[['Name', 'Title']].head())
# Create AgeGroup column: Bin ages into categories
bins = [0, 12, 18, 35, 60, 99]
labels = ['Child', 'Teen', 'Adult', 'Middle Aged', 'Senior']
df['AgeGroup'] = pd.cut(df['Age'], bins=bins, labels=labels)
print(df[['Age', 'AgeGroup']].head(10))
# Handle missing ages by filling with AgeGroup median
medians = df.groupby('AgeGroup')['Age'].transform('median')
df['AgeFilled'] = df['Age'].fillna(medians)
print(df.loc[df['Age'].isnull(), ['AgeGroup', 'Age', 'AgeFilled']].head())
# Create IsAlone feature: 1 if passenger is alone, else 0
df['IsAlone'] = (df['FamilySize'] == 1).astype(int)
print(df[['FamilySize', 'IsAlone']].head(8))
# Make a categorical variable: Deck extracted from Cabin
df['Deck'] = df['Cabin'].str[0]
print(df[['Cabin', 'Deck']].head(8))
# Create Fare per person: Fare divided by FamilySize
df['FarePerPerson'] = df['Fare'] / df['FamilySize']
print(df[['Fare', 'FamilySize', 'FarePerPerson']].head(8))
# Map Sex column to numeric for analysis
df['SexNum'] = df['Sex'].map({'male': 0, 'female': 1})
print(df[['Sex', 'SexNum']].head(8))
# One-hot encode Embarked column
embark_dummies = pd.get_dummies(df['Embarked'], prefix='Embarked')
df = pd.concat([df, embark_dummies], axis=1)
print(df[['Embarked', 'Embarked_C', 'Embarked_Q', 'Embarked_S']].head(8))
# Create interaction feature: Pclass by Sex
df['ClassSex'] = df['Pclass'].astype(str) + '_' + df['Sex']
print(df[['Pclass', 'Sex', 'ClassSex']].head(8))
# Detect outliers in Fare using a Z-score method
fare_mean = df['Fare'].mean()
fare_std = df['Fare'].std()
df['FareZscore'] = (df['Fare'] - fare_mean) / fare_std
print(df[['Fare', 'FareZscore']].head(8))
Mini-Project: Feature Engineering EDA#
Let us pause and practice turning our new features into insight.
We will ask if traveling alone or with family changed survival rates.
Try to guess before running the code below!
# Group by IsAlone and calculate mean survival
print(df.groupby('IsAlone')['Survived'].mean())
# Average survival by AgeGroup
print(df.groupby('AgeGroup')['Survived'].mean())
# Visualization: Survival by AgeGroup and Sex
import matplotlib.pyplot as plt
pd.crosstab(df['AgeGroup'], df['Sex']).plot(kind='bar', stacked=True, figsize=(6,4))
plt.title('Count of Passengers by Age Group and Sex')
plt.xlabel('Age Group')
plt.ylabel('Count')
plt.legend(title='Sex')
plt.tight_layout()
plt.show()
# Quick correlation check for new features
new_features = ['FamilySize', 'IsAlone', 'SexNum', 'FarePerPerson']
print(df[new_features + ['Survived']].corr())
# Caution: Avoid data leakage
df['SurvivalKnown'] = df['Survived'] # This would leak the target.
print('Columns:', df.columns.tolist())
del df['SurvivalKnown']
Recap: What We Learned#
- Creating features by combining columns
- Extracting information from text
- Handling missing values
- Making categories
- Encoding labels for models
- Checking for outliers and leakage
Practice these steps to improve any data project!
Keep Practicing & Subscribe!#
Want more pandas tutorials and tips?
Practice: Pick a fun dataset, design and test at least two of your own features.
Subscribe for future walkthroughs on pandas, data science, and machine learning!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



