Lesson 29 · Mastering Pandas
Master Grouping and Aggregation in Python for Effective Data Analysis
Grouping and aggregation let you answer powerful business questions. You will learn to use groupby(), calculate summaries, and uncover trends. We will…
- CourseMastering Pandas
- Lesson29 of 44
- Video14 min
- FormatJupyter notebook · 16 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbLesson Overview: Grouping Data and Aggregation with pandas#
Grouping and aggregation let you answer powerful business questions. You will learn to use groupby(), calculate summaries, and uncover trends.
We will practice with the Titanic dataset, a classic for passenger analytics!
import warnings
import numpy as np
np.random.seed(42)
warnings.filterwarnings("ignore")
# Data setup (Titanic Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
Exploring the Titanic Data#
Before we dive into grouping, notice columns like 'Pclass', 'Sex', and 'Survived'. These will help us segment the data in meaningful ways.
# Check for missing values
df.isnull().sum()
# Fill missing 'Age' values with the median
df['Age'].fillna(df['Age'].median(), inplace=True)
# See unique passenger classes
print(df['Pclass'].unique())
What is groupby in pandas?#
groupby() lets you split your data into groups based on the values in one or more columns. For example, you can ask: what is the average age of survivors in each class?
# Group by 'Pclass' and get mean age
pclass_age = df.groupby('Pclass')['Age'].mean()
print(pclass_age)
# Group by 'Sex' and see survival rate
survival_rate = df.groupby('Sex')['Survived'].mean()
print(survival_rate)
# Group by multiple columns: class and sex
survival_by_group = df.groupby(['Pclass', 'Sex'])['Survived'].mean()
print(survival_by_group)
# Find mean, median, and max fare by class
fare_stats = df.groupby('Pclass')['Fare'].agg(['mean', 'median', 'max'])
print(fare_stats)
Quick practice: Group and summarize#
Try grouping by 'Embarked' and counting how many passengers from each port survived.
# Group by 'Embarked' and 'Pclass', counting survivors per group
survivor_counts = df.groupby(['Embarked', 'Pclass'])['Survived'].sum().unstack()
print(survivor_counts)
# Group by 'Sex' and describe ages
age_stats = df.groupby('Sex')['Age'].describe()
print(age_stats)
# Custom aggregation: average fare per survival status and class
def mean_fare(x):
return round(x.mean(), 2)
fare_grouped = df.groupby(['Survived', 'Pclass'])['Fare'].agg(mean_fare)
print(fare_grouped)
# Filter: Only show groups with more than 100 passengers
sizes = df.groupby('Pclass').filter(lambda x: len(x) > 100)
print(sizes['Pclass'].value_counts())
# Pivot table recreation: survival rate by class and gender
pivot = df.pivot_table(values='Survived', index='Pclass', columns='Sex', aggfunc='mean')
print(pivot)
# Visualize groupby: average age by class
import matplotlib.pyplot as plt
grouped_avg_age = df.groupby('Pclass')['Age'].mean()
grouped_avg_age.plot(kind='bar', color='skyblue')
plt.ylabel('Average Age')
plt.title('Average Age by Passenger Class')
plt.show()
Challenge: Find the youngest and oldest survivor in each class#
Use groupby and aggregation to figure this out!
# Youngest and oldest survivor per class
survivors = df[df['Survived'] == 1]
result = survivors.groupby('Pclass')['Age'].agg(['min', 'max'])
print(result)
Recap: Groupby and Aggregations in pandas#
- groupby lets you split, summarize, and re-combine data.
- Aggregations provide flexible summaries like mean, min, max, and more.
- Try groupby with your own datasets and real-world questions!
Want more data science tutorials?#
Subscribe to our channel for more pandas videos and request topics in the comments. Your feedback shapes our next deep dive!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



