Mathew K Analytics

Lesson 23 · Mastering Pandas

Mastering Pandas DataFrames: Applying Functions Across Rows & Columns for Data Analysis

Have you ever wanted to quickly transform or summarize columns in your DataFrame? In this lesson, we learn how to apply functionsboth built-in and…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Applying Functions Across Columns and Rows in Pandas#

Have you ever wanted to quickly transform or summarize columns in your DataFrame?

In this lesson, we learn how to apply functionsboth built-in and customacross columns or rows in pandas.

This skill unlocks serious power for data cleaning, feature engineering, and flexible analysis.

Let us get started with a practical example using the Titanic dataset.

# Data setup (Titanic Dataset)
import warnings; warnings.filterwarnings('ignore')
import pandas as pd
import numpy as np
np.random.seed(42)
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
(891, 12)
   PassengerId  Survived  Pclass  \
0            1         0       3   
1            2         1       1   
2            3         1       3   

                                                Name     Sex   Age  SibSp  \
0                            Braund, Mr. Owen Harris    male  22.0      1   
1  Cumings, Mrs. John Bradley (Florence Briggs Th...  female  38.0      1   
2                             Heikkinen, Miss. Laina  female  26.0      0   

   Parch            Ticket     Fare Cabin Embarked  
0      0         A/5 21171   7.2500   NaN        S  
1      0          PC 17599  71.2833   C85        C  
2      0  STON/O2. 3101282   7.9250   NaN        S  

Why apply functions?#

Applying functions helps us:

  • Quickly create new columns based on old ones
  • Standardize or transform data
  • Summarize row or column information

Let us see how to use the pandas apply method.

# Checking columns in our Titanic DataFrame
print(df.columns.tolist())
['PassengerId', 'Survived', 'Pclass', 'Name', 'Sex', 'Age', 'SibSp', 'Parch', 'Ticket', 'Fare', 'Cabin', 'Embarked']
# Applying a built-in function to a column: get the length of each passenger's name
df['name_length'] = df['Name'].apply(len)
df[['Name', 'name_length']].head(3)
Name name_length
0 Braund, Mr. Owen Harris 23
1 Cumings, Mrs. John Bradley (Florence Briggs Th... 51
2 Heikkinen, Miss. Laina 22
# Using apply with a custom function: check if a passenger is a child (under 16)
def is_child(age):
    if pd.isnull(age):
        return None
    return age < 16

df['is_child'] = df['Age'].apply(is_child)
df[['Name', 'Age', 'is_child']].head(5)
Name Age is_child
0 Braund, Mr. Owen Harris 22.0 False
1 Cumings, Mrs. John Bradley (Florence Briggs Th... 38.0 False
2 Heikkinen, Miss. Laina 26.0 False
3 Futrelle, Mrs. Jacques Heath (Lily May Peel) 35.0 False
4 Allen, Mr. William Henry 35.0 False

The power of 'apply' with lambda functions#

Lambdas are small, anonymous functions.

They let us write quick logic directly inside 'apply'.

Let us see how we can create a new column on the fly.

# Quickly check who paid over $50 for a ticket using lambda
df['BigSpender'] = df['Fare'].apply(lambda x: x > 50 if pd.notnull(x) else False)
df[['Fare', 'BigSpender']].head(8)
Fare BigSpender
0 7.2500 False
1 71.2833 True
2 7.9250 False
3 53.1000 True
4 8.0500 False
5 8.4583 False
6 51.8625 True
7 21.0750 False
# Apply a function across several columns using axis=1
def family_size(row):
    return row['SibSp'] + row['Parch'] + 1  # Include self

df['FamilySize'] = df.apply(family_size, axis=1)
df[['Name', 'SibSp', 'Parch', 'FamilySize']].head(5)
Name SibSp Parch FamilySize
0 Braund, Mr. Owen Harris 1 0 2
1 Cumings, Mrs. John Bradley (Florence Briggs Th... 1 0 2
2 Heikkinen, Miss. Laina 0 0 1
3 Futrelle, Mrs. Jacques Heath (Lily May Peel) 1 0 2
4 Allen, Mr. William Henry 0 0 1
# Using applymap: Apply a function to every element in a DataFrame
df2 = df[['SibSp', 'Parch']].copy()
df2_applied = df2.applymap(lambda x: x * 10)
print(df2_applied.head())
   SibSp  Parch
0     10      0
1     10      0
2      0      0
3     10      0
4      0      0

Practice: Try describing each passenger#

You can use apply to join columnstry creating a 'Description' column that combines name, age, and sex.

Think about when to use 'apply', 'applymap', or just vectorized operations!

Next, let us look at two special pandas tricks: 'agg' and 'transform'.

# Using agg to summarize with multiple functions
df[['Fare', 'Age']].agg(['min', 'max', 'mean'])
Fare Age
min 0.000000 0.420000
max 512.329200 80.000000
mean 32.204208 29.699118
# Using transform to scale the Fare column
fare_mean = df['Fare'].mean()
fare_std = df['Fare'].std()
df['Fare_zscore'] = df['Fare'].transform(lambda x: (x - fare_mean) / fare_std)
df[['Fare', 'Fare_zscore']].head(5)
Fare Fare_zscore
0 7.2500 -0.502163
1 71.2833 0.786404
2 7.9250 -0.488580
3 53.1000 0.420494
4 8.0500 -0.486064
# Combining apply with filtering: find the longest name among survivors
survivors = df[df['Survived'] == 1]
longest_name = survivors['Name'].apply(len).max()
name_row = survivors[survivors['Name'].apply(len) == longest_name][['Name', 'Survived']]
print(name_row)
                                                  Name  Survived
307  Penasco y Castellana, Mrs. Victor de Satode (M...         1
# Challenge: Prompt user to type a minimum fare and flag passengers above it
min_fare = float(input('Enter a minimum fare: '))
df['BigSpenderCustom'] = df['Fare'].apply(lambda f: f > min_fare if pd.notnull(f) else False)
print(df[['Fare', 'BigSpenderCustom']].head())
      Fare  BigSpenderCustom
0   7.2500             False
1  71.2833              True
2   7.9250             False
3  53.1000              True
4   8.0500             False
# Real-world use: flag cabin letters (A, B, etc.) with apply and lambda
def cabin_letter(cabin):
    if pd.isnull(cabin):
        return None
    return str(cabin)[0]

df['CabinLetter'] = df['Cabin'].apply(cabin_letter)
df[['Cabin', 'CabinLetter']].head(8)
Cabin CabinLetter
0 NaN None
1 C85 C
2 NaN None
3 C123 C
4 NaN None
5 NaN None
6 E46 E
7 NaN None

Troubleshooting: Why does my apply code break?#

Common issues:

  • Your function expects a column but you pass a row (or the other way around).
  • You forget to set axis=1 when looping by rows.
  • You do not handle missing values safely.
  • Your lambda has a typo or missing colon.
  • You apply to an object column, but want a number.

Always check error messages and preview outputs using .head().

# Common error demonstration: applying a function the wrong way
try:
    df.apply(len)
except Exception as e:
    print('Caught:', e)
    
# Best Practice: Use vectorized pandas methods when possible
df['name_length_vec'] = df['Name'].str.len()
df[['Name', 'name_length_vec', 'name_length']].head(3)
Name name_length_vec name_length
0 Braund, Mr. Owen Harris 23 23
1 Cumings, Mrs. John Bradley (Florence Briggs Th... 51 51
2 Heikkinen, Miss. Laina 22 22

Mini-Project: Feature Creation for Survival Prediction#

Suppose we want to analyze how custom features relate to survival. We build and inspect a couple new columns using apply and basic logic.

Let us see which features separate survivors from non-survivors.

# Who survived with a large family? Using our FamilySize feature
print(df.groupby('Survived')['FamilySize'].mean())
Survived
0    1.883424
1    1.938596
Name: FamilySize, dtype: float64
# Combining multiple apply-created features in a scatter plot
import matplotlib.pyplot as plt
plt.figure(figsize=(7,4))
plt.scatter(df['FamilySize'], df['Fare'], c=df['Survived'], cmap='coolwarm', alpha=0.5)
plt.xlabel('Family Size')
plt.ylabel('Fare')
plt.title('Fare vs Family Size (Color: Survived)')
plt.show()
No description has been provided for this image

Recap: Using apply, applymap, agg, and transform#

Today you learned to:

  • Apply built-in and custom functions to columns
  • Create new features by row or column
  • Use lambdas for quick logic
  • Aggregate and transform data for reporting and modeling

Try combining these tools in your next analysis!

Thanks for joining this episode!

Practice by writing your own mini-features on a favorite dataset.

Tell us which pandas topics you want to see next.

If you found this helpful, please subscribe and share!

See you in the next lesson.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.