Mathew K Analytics

Lesson 24 · Mastering Pandas

Master Lambda Functions & Vectorized Operations in Pandas for Faster Data Analysis

In this lesson, we dig into lambda functions and powerful vectorized operations. Ever wondered how to apply custom calculations efficiently to your data?…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Lambda Functions and Vectorized Operations in Pandas#

In this lesson, we dig into lambda functions and powerful vectorized operations.

Ever wondered how to apply custom calculations efficiently to your data?

You will learn ways to write functions that run super-fast over large tables, without any loops.

These skills help you wrangle, clean, and analyze data like a pro.

Ready? Let us get started!

import warnings
import numpy as np
np.random.seed(42)
warnings.filterwarnings("ignore")
 
 
# Data setup (Restaurant Tips Dataset)
import seaborn as sns
df = sns.load_dataset('tips')
print(df.shape)
print(df.head(3))
(244, 7)
   total_bill   tip     sex smoker  day    time  size
0       16.99  1.01  Female     No  Sun  Dinner     2
1       10.34  1.66    Male     No  Sun  Dinner     3
2       21.01  3.50    Male     No  Sun  Dinner     3

Why use lambda functions and vectorized operations?#

Pandas works best when you avoid loops.

Vectorized code runs on whole columns at once. Lambda functions let you quickly write new calculations.

They are key for clean, efficient data work.

# Vectorized math: add a "total" column
df['total'] = df['total_bill'] + df['tip']
print(df[['total_bill','tip','total']].head(3))
   total_bill   tip  total
0       16.99  1.01  18.00
1       10.34  1.66  12.00
2       21.01  3.50  24.51
# Vectorized math with conditions
df['big_tip'] = df['tip'] > 5
print(df[['tip','big_tip']].head())
    tip  big_tip
0  1.01    False
1  1.66    False
2  3.50    False
3  3.31    False
4  3.61    False

What are lambda functions?#

A lambda function is a short, one-line function.

In pandas, we often use lambdas inside apply().

Syntax example: lambda x: x + 5

# Using lambda with apply() to tip column
df['tip_twice'] = df['tip'].apply(lambda x: x * 2)
print(df[['tip','tip_twice']].head())
    tip  tip_twice
0  1.01       2.02
1  1.66       3.32
2  3.50       7.00
3  3.31       6.62
4  3.61       7.22
# Conditional logic inside lambda
df['service_label'] = df['tip'].apply(lambda x: 'generous' if x > 5 else 'normal')
print(df[['tip','service_label']].head(6))
    tip service_label
0  1.01        normal
1  1.66        normal
2  3.50        normal
3  3.31        normal
4  3.61        normal
5  4.71        normal
# Lambda with multiple columns via apply (axis=1)
df['bill_ratio'] = df.apply(lambda row: row['tip'] / row['total_bill'], axis=1)
print(df[['total_bill','tip','bill_ratio']].head(4))
   total_bill   tip  bill_ratio
0       16.99  1.01    0.059447
1       10.34  1.66    0.160542
2       21.01  3.50    0.166587
3       23.68  3.31    0.139780

Vectorized vs. apply: Which is faster?#

  • Vectorized math is much faster: runs in C under the hood.
  • apply() and lambda are more flexible but slower for big data.

Tip: Prefer vectorized solutions when possible for speed.

# Vectorized string operation: add "$" to tip column
df['tip_str'] = df['tip'].astype(str) + " $"
print(df[['tip','tip_str']].head(5))
    tip tip_str
0  1.01  1.01 $
1  1.66  1.66 $
2  3.50   3.5 $
3  3.31  3.31 $
4  3.61  3.61 $
# Using .map() with a function
def bonus(size):
        if size >= 5:
                        return 1
        else:
                        return 0
df['party_bonus'] = df['size'].map(bonus)
print(df[['size','party_bonus']].head(6))
   size  party_bonus
0     2            0
1     3            0
2     3            0
3     2            0
4     4            0
5     4            0
# Using np.where() for vectorized conditional assignment
import numpy as np
df['meal_value'] = np.where(df['total_bill'] > 30, 'expensive', 'standard')
print(df[['total_bill','meal_value']].head(7))
   total_bill meal_value
0       16.99   standard
1       10.34   standard
2       21.01   standard
3       23.68   standard
4       24.59   standard
5       25.29   standard
6        8.77   standard

Mini-project: Calculate tip percent and categorize#

Challenge: Create two new columns:

  1. tip_percent: (tip divided by total_bill) times 100, rounded to one decimal place.
  2. tip_level: Use a lambda with apply to label as 'high' if tip_percent >= 20 else 'normal'.

Ready? Let us code this together.

# Create tip_percent column (vectorized)
df['tip_percent'] = (df['tip'] / df['total_bill'] * 100).round(1)
print(df[['tip','total_bill','tip_percent']].head(8))
    tip  total_bill  tip_percent
0  1.01       16.99          5.9
1  1.66       10.34         16.1
2  3.50       21.01         16.7
3  3.31       23.68         14.0
4  3.61       24.59         14.7
5  4.71       25.29         18.6
6  2.00        8.77         22.8
7  3.12       26.88         11.6
# Create tip_level column (lambda + apply)
df['tip_level'] = df['tip_percent'].apply(lambda x: 'high' if x >= 20 else 'normal')
print(df[['tip_percent','tip_level']].head(10))
   tip_percent tip_level
0          5.9    normal
1         16.1    normal
2         16.7    normal
3         14.0    normal
4         14.7    normal
5         18.6    normal
6         22.8      high
7         11.6    normal
8         13.0    normal
9         21.9      high
# Vectorized math with missing values
df_missing = df.copy()
df_missing.loc[0,'tip'] = np.nan  # Set one tip to NaN
df_missing['tip_plus_five'] = df_missing['tip'] + 5
print(df_missing[['tip','tip_plus_five']].head(3))
    tip  tip_plus_five
0   NaN            NaN
1  1.66           6.66
2  3.50           8.50
# Clean missing values with .fillna() and re-apply math
df_missing['tip_filled'] = df_missing['tip'].fillna(0)
df_missing['new_total'] = df_missing['tip_filled'] + df_missing['total_bill']
print(df_missing[['tip','tip_filled','new_total']].head(4))
    tip  tip_filled  new_total
0   NaN        0.00      16.99
1  1.66        1.66      12.00
2  3.50        3.50      24.51
3  3.31        3.31      26.99

Troubleshooting common lambda and vectorization errors#

Common issues:

  • Forgetting axis=1 for row-based apply.
  • Mismatched data types.
  • Using a string method on a numeric column.
  • Not handling missing values.

Pandas usually gives a helpful error message. Check column types or try df.info() to debug.

# Challenge practice: flag rows where tip is above mean
mean_tip = df['tip'].mean()
df['above_average_tip'] = df['tip'] > mean_tip
print(f"Average tip: {mean_tip:.2f}")
print(df[['tip','above_average_tip']].sample(5, random_state=42))
Average tip: 3.00
      tip  above_average_tip
24   3.18               True
6    2.00              False
153  2.00              False
211  5.16               True
198  2.00              False

Recap: What did you learn?#

  • Lambda functions let you quickly apply logic to columns.
  • Vectorized math is super fast and clean.
  • Use apply(axis=1) for row-based calculations with several columns.
  • Handle missing values and check types to avoid bugs.

Practice these skills with your own data for speed and clarity.

Want more data skills? Subscribe and check out our next pandas lesson on groupby tricks and real-world analysis!#

Thanks for learning with us!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.