Lesson 8 · Mastering Pandas
Exploring and Cleaning Real-World Datasets with Python Pandas: A Practical Guide
Welcome! In this lesson, you will learn how to use Python pandas to explore and analyze real-world datasets. By the end, you will be comfortable: loading…
- CourseMastering Pandas
- Lesson8 of 44
- Video19 min
- FormatJupyter notebook · 15 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbExploring Real Datasets with Pandas#
Welcome! In this lesson, you will learn how to use Python pandas to explore and analyze real-world datasets.
By the end, you will be comfortable: loading data, cleaning data, filtering, grouping, joining, reshaping, visualizing, and building mini-projects.
Let us get started!
import warnings; warnings.filterwarnings("ignore")
import numpy as np
np.random.seed(42)
# Suppress all warnings so your notebook stays clean
Step 1: Loading the Titanic Dataset#
We will use the Titanic dataset, a classic for data exploration. It captures information about passengers, such as age, gender, fare, and survival.
Let us load and preview this dataset.
import pandas as pd
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
Step 2: Basic Exploration#
Let us quickly check columns, data types, and missing values.
This helps us understand the shape and health of our data.
print(df.columns)
print(df.dtypes)
print(df.isnull().sum())
Step 3: Quick Statistics#
Pandas can describe numerical and text columns for you. Check means, counts, and more with a single function.
Let us see a summary.
print(df.describe(include='all'))
Step 4: Cleaning Data#
Messy data is very common. Let us start simple: fill missing age values.
We will fill them with the average age.
mean_age = df['Age'].mean()
df['Age'].fillna(mean_age, inplace=True)
print(df['Age'].isnull().sum())
Step 5: Selecting and Filtering#
Let us look for passengers who paid high fares.
Filtering helps you focus on interesting data.
rich = df[df['Fare'] > 100]
print(rich[['Name', 'Fare']].head())
Step 6: Grouping and Aggregation#
Pandas makes it easy to group data and calculate statistics by group.
Let us find the survival rate by passenger class.
grouped = df.groupby('Pclass')['Survived'].mean()
print(grouped)
Step 7: Joining DataFrames#
You often work with several tables of data. Let us join two DataFrames together on a shared column.
We will make a little demo passengers table and join on the Name.
demo_passengers = pd.DataFrame({
'Name': [df.iloc[0]['Name'], df.iloc[1]['Name']],
'Extra_Info': ['Friend', 'Colleague']
})
merged = pd.merge(df, demo_passengers, on='Name', how='left')
print(merged[['Name', 'Extra_Info']].head(3))
Step 8: Pivot Tables and Reshaping#
Pivot tables reshape your data, making summaries by groups.
Let us see the average age by sex and class.
pivot = df.pivot_table(index='Sex', columns='Pclass', values='Age', aggfunc='mean')
print(pivot)
Step 9: Time-Series Handling#
Suppose our Titanic data included a Date column. For now, let us use a real time-series example: the Flights dataset.
We will load it and plot monthly airline passengers over time.
import seaborn as sns
flights_df = sns.load_dataset('flights')
print(flights_df.shape)
print(flights_df.head(3))
import matplotlib.pyplot as plt
monthly = flights_df.pivot_table(index='year', columns='month', values='passengers')
flights_df['date'] = pd.to_datetime(flights_df['year'].astype(str) + '-' + flights_df['month'].astype(str) + '-01')
plt.figure(figsize=(8,4))
plt.plot(flights_df['date'], flights_df['passengers'])
plt.xlabel('Date')
plt.ylabel('Passengers')
plt.title('Monthly Airline Passengers Over Time')
plt.tight_layout()
plt.show()
Step 10: Visualization with Pandas#
Pandas can also plot data directly.
Let us visualize the number of survivors by sex from the Titanic dataset.
df.groupby('Sex')['Survived'].sum().plot(kind='bar', color=['skyblue', 'lightcoral'])
plt.ylabel('Survivors')
plt.title('Survivors by Sex')
plt.show()
Step 11: Mini-Project Part 1 Exploratory Data Analysis (Titanic)#
Now let us do a hands-on mini project.
We will answer: What are some factors linked with survival? Which age group survived most?
Let us explore age, class, and survival together.
df['AgeGroup'] = pd.cut(df['Age'], bins=[0,12,18,40,80], labels=['Child','Teen','Adult','Senior'])
grouped_survival = df.groupby(['AgeGroup','Pclass'])['Survived'].mean().unstack()
print(grouped_survival)
# Visualize survival rates by age group and class
grouped_survival.plot(kind='bar', figsize=(8,5))
plt.ylabel('Survival Rate')
plt.title('Survival Rate by Age Group and Class')
plt.show()
Step 12: Mini-Project Part 2 Feature Engineering#
Let us try creating a new feature: FamilySize. Then, we can check how family size relates to survival.
This is common when preparing data for machine learning.
df['FamilySize'] = df['SibSp'] + df['Parch'] + 1
print(df[['SibSp', 'Parch', 'FamilySize']].head())
family_survival = df.groupby('FamilySize')['Survived'].mean()
print(family_survival)
Step 13: Best Practices and Performance Tips#
- Always check for missing or strange data right away.
- Use .copy() if you want a real copy of a DataFrame, not a reference.
- For large datasets, try .info(), .head(), and .sample() to save time.
- Chain commands in separate steps for clarity.
- Use vectorized operations over loops; they are much faster in pandas.
Step 14: Troubleshooting Common Errors#
Common problems:
- KeyError: Check for typos in column names.
- ValueError when reshaping: Watch your index and column labels.
- SettingWithCopyWarning: Use .loc or .copy() to avoid silent bugs.
If stuck, try printing shapes and dtypes, or use df.sample(5) for a quick sanity check.
try:
print(df['NotARealColumn'].head())
except KeyError as e:
print(f"Error: {e}")
Step 15: Challenge Exercises#
- Can you find the top 3 oldest survivors?
- Make a new column for fare per person (Fare divided by FamilySize).
- Explore another public dataset and produce a simple pivot table.
Pause the video and give these a try!
Recap#
Congratulations on completing this lesson!
You have loaded, explored, cleaned, filtered, grouped, reshaped, and visualized real-world data. You are ready to start your own projects.
Keep practicing and experimenting.
Want more tutorials?#
Subscribe to the channel and leave a comment with your favorite dataset or question you want answered next!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



