Lesson 51 · Mastering Pandas
Step-by-Step Guide to Exploratory Data Analysis (EDA) Using Python and Pandas
Welcome! In this lesson, we will explore real data and master pandas for EDA. We will analyze the world-famous Titanic dataset to find insights for survival…
- CourseMastering Pandas
- Lesson51 of 44
- Video27 min
- FormatJupyter notebook · 22 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbHands-On Exploratory Data Analysis (EDA) Project in Pandas#
Welcome! In this lesson, we will explore real data and master pandas for EDA.
We will analyze the world-famous Titanic dataset to find insights for survival analysis and data-driven storytelling.
You will learn data loading, cleaning, exploration, aggregation, and more!
Lets get started.
import warnings
warnings.filterwarnings("ignore")
# Data setup (Titanic Dataset)
import pandas as pd
import numpy as np
np.random.seed(42)
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
Understanding the Data Structure#
Lets quickly learn about the columns or features in our Titanic dataset.
Each row tells us about a unique passenger and their attributes, such as:
- PassengerId
- Survived (1=Yes, 0=No)
- Pclass (ticket class)
- Name
- Sex
- Age
- SibSp (siblings/spouses aboard)
- Parch (parents/children aboard)
- Ticket
- Fare
- Cabin
- Embarked (port of embarkation)
# Look at column names and data types
print(df.dtypes)
# Basic data info: missing values, non-null counts
print(df.info())
# Checking for missing values directly
print(df.isnull().sum())
Quick Data Cleaning Steps#
Handling missing values is a core EDA task.
Lets decide what to do with them, especially for major columns like 'Age' and 'Cabin'.
Option 1: Remove rows or Option 2: Fill in with statistics (like median age).
Lets clean the 'Age' column by filling missing values with the median.
# Fill missing Age with median age
median_age = df['Age'].median()
df['Age'] = df['Age'].fillna(median_age)
print('Missing Age values:', df['Age'].isnull().sum())
# Fill missing 'Embarked' (port) with the most common port
most_common_port = df['Embarked'].mode()[0]
df['Embarked'] = df['Embarked'].fillna(most_common_port)
print('Missing Embarked values:', df['Embarked'].isnull().sum())
# Drop the 'Cabin' column (too many missing values)
df = df.drop('Cabin', axis=1)
print(df.columns)
Exploring Patterns: Summary Statistics#
Summary statistics help us quickly understand the range of values in our data.
Lets look at the basics like mean, median, min, max for numeric columns.
# Describe numeric columns
print(df.describe())
# What percent survived overall?
survival_rate = df['Survived'].mean() * 100
print(f"Survival rate: {survival_rate:.2f}%")
# Survival rate by gender
print(df.groupby('Sex')['Survived'].mean())
# Survival rate by passenger class
print(df.groupby('Pclass')['Survived'].mean())
Filtering and Conditional Selection#
Lets learn to pick rows that match certain conditions, like all children, or people who embarked from a specific port.
This skill is essential for quick hypothesis testing.
# Show all passengers under age 12 (children)
children = df[df['Age'] < 12]
print(children[['Name', 'Age', 'Survived']].head())
# Passengers who paid high fares (above 100)
rich = df[df['Fare'] > 100]
print(rich[['Name', 'Fare', 'Survived']].head())
Aggregation and GroupBy: Average Age by Class#
Lets discover more by grouping passengers and computing aggregate statistics.
We will find the typical age and survival chance for each class.
# Average age and survival by class
print(df.groupby('Pclass')[['Age', 'Survived']].mean())
# Number of survivors by port of embarkation
print(df.groupby('Embarked')['Survived'].sum())
Pivot Tables: Survival by Class and Gender#
A pivot table can show patterns across two or more variables at once.
Lets create a table to spot differences in survival rates by both class and gender.
# Create a pivot table of survival rates by class and sex
pivot = df.pivot_table(index='Pclass', columns='Sex', values='Survived', aggfunc='mean')
print(pivot)
Simple Data Visualization: Survival by Age#
Visualizations help us see trends and relationships that numbers alone may hide.
Lets plot a histogram of ages for survivors and non-survivors side by side.
import matplotlib.pyplot as plt
plt.figure(figsize=(8,4))
df[df['Survived']==1]['Age'].hist(alpha=0.6, bins=20, label='Survived')
df[df['Survived']==0]['Age'].hist(alpha=0.6, bins=20, label='Did not survive')
plt.legend()
plt.xlabel('Age')
plt.ylabel('Number of passengers')
plt.title('Age Distribution by Survival')
plt.show()
Mini Project: Your Turn#
Lets combine your new skills into a simple EDA 'mini-project' question:
Question: Did traveling alone or with family make a difference for survival?
You can use the 'SibSp' and 'Parch' columns to define if a person was alone (both zero) or with family (either above zero).
Lets try to answer this step by step!
# Add a column 'Alone' (True if no siblings/spouses and no parents/children)
df['Alone'] = (df['SibSp'] == 0) & (df['Parch'] == 0)
print(df[['SibSp', 'Parch', 'Alone']].head())
# Compare survival for people alone vs with family
print(df.groupby('Alone')['Survived'].mean())
Common Problems in Pandas and How to Fix Them#
- Typos in column names can cause errors.
- Writing 'df[Age]' instead of 'df["Age"]' looks for a variable, not a column.
- Chaining many methods at once can sometimes return a copy, not the real DataFrame.
If something is not working, check your column names and try printing out interim results.
# Practice: Try renaming a column
df = df.rename(columns={'Fare': 'TicketFare'})
print(df.columns)
# Practice: Try summarizing with value_counts
print(df['Embarked'].value_counts())
# Quick user challenge: input() for a guessHow many passengers traveled solo?
guess = int(input("Guess: How many Titanic passengers were alone? "))
answer = df['Alone'].sum()
print(f"You guessed: {guess}. The actual answer is {answer}!")
Review and Next Steps#
Great job! You loaded, cleaned, explored, grouped, visualized, and asked questions about real Titanic data.
Keep practicing: Try similar steps with new datasets, or repeat with more questions of your own.
Thanks for joining us. For more learning, subscribe for more videos and try out the code on your own!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



