Lesson 23 · Mastering Pandas
Mastering Pandas DataFrames: Applying Functions Across Rows & Columns for Data Analysis
Have you ever wanted to quickly transform or summarize columns in your DataFrame? In this lesson, we learn how to apply functionsboth built-in and…
- CourseMastering Pandas
- Lesson23 of 44
- Video12 min
- FormatJupyter notebook · 16 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbApplying Functions Across Columns and Rows in Pandas#
Have you ever wanted to quickly transform or summarize columns in your DataFrame?
In this lesson, we learn how to apply functionsboth built-in and customacross columns or rows in pandas.
This skill unlocks serious power for data cleaning, feature engineering, and flexible analysis.
Let us get started with a practical example using the Titanic dataset.
# Data setup (Titanic Dataset)
import warnings; warnings.filterwarnings('ignore')
import pandas as pd
import numpy as np
np.random.seed(42)
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
Why apply functions?#
Applying functions helps us:
- Quickly create new columns based on old ones
- Standardize or transform data
- Summarize row or column information
Let us see how to use the pandas apply method.
# Checking columns in our Titanic DataFrame
print(df.columns.tolist())
# Applying a built-in function to a column: get the length of each passenger's name
df['name_length'] = df['Name'].apply(len)
df[['Name', 'name_length']].head(3)
# Using apply with a custom function: check if a passenger is a child (under 16)
def is_child(age):
if pd.isnull(age):
return None
return age < 16
df['is_child'] = df['Age'].apply(is_child)
df[['Name', 'Age', 'is_child']].head(5)
The power of 'apply' with lambda functions#
Lambdas are small, anonymous functions.
They let us write quick logic directly inside 'apply'.
Let us see how we can create a new column on the fly.
# Quickly check who paid over $50 for a ticket using lambda
df['BigSpender'] = df['Fare'].apply(lambda x: x > 50 if pd.notnull(x) else False)
df[['Fare', 'BigSpender']].head(8)
# Apply a function across several columns using axis=1
def family_size(row):
return row['SibSp'] + row['Parch'] + 1 # Include self
df['FamilySize'] = df.apply(family_size, axis=1)
df[['Name', 'SibSp', 'Parch', 'FamilySize']].head(5)
# Using applymap: Apply a function to every element in a DataFrame
df2 = df[['SibSp', 'Parch']].copy()
df2_applied = df2.applymap(lambda x: x * 10)
print(df2_applied.head())
Practice: Try describing each passenger#
You can use apply to join columnstry creating a 'Description' column that combines name, age, and sex.
Think about when to use 'apply', 'applymap', or just vectorized operations!
Next, let us look at two special pandas tricks: 'agg' and 'transform'.
# Using agg to summarize with multiple functions
df[['Fare', 'Age']].agg(['min', 'max', 'mean'])
# Using transform to scale the Fare column
fare_mean = df['Fare'].mean()
fare_std = df['Fare'].std()
df['Fare_zscore'] = df['Fare'].transform(lambda x: (x - fare_mean) / fare_std)
df[['Fare', 'Fare_zscore']].head(5)
# Combining apply with filtering: find the longest name among survivors
survivors = df[df['Survived'] == 1]
longest_name = survivors['Name'].apply(len).max()
name_row = survivors[survivors['Name'].apply(len) == longest_name][['Name', 'Survived']]
print(name_row)
# Challenge: Prompt user to type a minimum fare and flag passengers above it
min_fare = float(input('Enter a minimum fare: '))
df['BigSpenderCustom'] = df['Fare'].apply(lambda f: f > min_fare if pd.notnull(f) else False)
print(df[['Fare', 'BigSpenderCustom']].head())
# Real-world use: flag cabin letters (A, B, etc.) with apply and lambda
def cabin_letter(cabin):
if pd.isnull(cabin):
return None
return str(cabin)[0]
df['CabinLetter'] = df['Cabin'].apply(cabin_letter)
df[['Cabin', 'CabinLetter']].head(8)
Troubleshooting: Why does my apply code break?#
Common issues:
- Your function expects a column but you pass a row (or the other way around).
- You forget to set
axis=1when looping by rows. - You do not handle missing values safely.
- Your lambda has a typo or missing colon.
- You apply to an object column, but want a number.
Always check error messages and preview outputs using .head().
# Common error demonstration: applying a function the wrong way
try:
df.apply(len)
except Exception as e:
print('Caught:', e)
# Best Practice: Use vectorized pandas methods when possible
df['name_length_vec'] = df['Name'].str.len()
df[['Name', 'name_length_vec', 'name_length']].head(3)
Mini-Project: Feature Creation for Survival Prediction#
Suppose we want to analyze how custom features relate to survival. We build and inspect a couple new columns using apply and basic logic.
Let us see which features separate survivors from non-survivors.
# Who survived with a large family? Using our FamilySize feature
print(df.groupby('Survived')['FamilySize'].mean())
# Combining multiple apply-created features in a scatter plot
import matplotlib.pyplot as plt
plt.figure(figsize=(7,4))
plt.scatter(df['FamilySize'], df['Fare'], c=df['Survived'], cmap='coolwarm', alpha=0.5)
plt.xlabel('Family Size')
plt.ylabel('Fare')
plt.title('Fare vs Family Size (Color: Survived)')
plt.show()
Recap: Using apply, applymap, agg, and transform#
Today you learned to:
- Apply built-in and custom functions to columns
- Create new features by row or column
- Use lambdas for quick logic
- Aggregate and transform data for reporting and modeling
Try combining these tools in your next analysis!
Thanks for joining this episode!
Practice by writing your own mini-features on a favorite dataset.
Tell us which pandas topics you want to see next.
If you found this helpful, please subscribe and share!
See you in the next lesson.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



