Lesson 69 · Mastering Pandas
Comprehensive Pandas Tutorial in Python: Data Analysis from Beginner to Advanced
Welcome to this hands-on Pandas tutorial! We will start from the basics and work our way through real-world data preparation, analysis, and visualization.…
- CourseMastering Pandas
- Lesson69 of 44
- Video36 min
- FormatJupyter notebook · 33 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbMastering Pandas: From Basics to Advanced#
Welcome to this hands-on Pandas tutorial! We will start from the basics and work our way through real-world data preparation, analysis, and visualization.
We will use the Titanic and Tips datasets. Both datasets come with Seaborn, so there is nothing extra to download.
You will learn how to explore, clean, transform, combine, and visualize data using Pandas.
Let us begin our journey!
Section 1: Data Setup and Exploration#
Let us load our first dataset (Titanic).
import seaborn as sns
import pandas as pd
import warnings
import numpy as np
np.random.seed(42)
warnings.filterwarnings("ignore")
# Load the Titanic dataset
df = sns.load_dataset("titanic")
# Preview the dataset
print("Shape:", df.shape)
print(df.head(3))
Exploring the DataFrame Structure#
Let us check the column names and data types in our dataset.
print("Columns:", df.columns.tolist())
print("\nData types:")
print(df.dtypes)
# How many unique values are in the "embarked" column?
print(df["embarked"].unique())
# Check basic statistics for numeric columns
print(df.describe())
Section 2: Data Cleaning and Preprocessing#
It is important to clean your data before analysis.
We will drop missing values for some basic tasks.
# Drop rows with any missing values
df_clean = df.dropna()
# Preview the cleaned data
print("Original shape:", df.shape)
print("Cleaned shape:", df_clean.shape)
# Fill missing 'age' values with the median
df_filled = df.copy()
median_age = df_filled['age'].median()
df_filled['age'] = df_filled['age'].fillna(median_age)
print(df_filled['age'].isnull().sum())
Checking for Duplicates#
Duplicate rows can cause problems in analysis.
Let us see if there are any duplicates.
print("Number of duplicate rows:", df.duplicated().sum())
Section 3: Filtering and Conditional Selection#
Let us work with the Tips dataset and practice selecting some rows.
import seaborn as sns
import pandas as pd
import warnings
warnings.filterwarnings("ignore")
# Load the Tips dataset
tips = sns.load_dataset("tips")
# Preview the dataset
print("Shape:", tips.shape)
print(tips.head(3))
# Select only the records where the tip is greater than $5
big_tips = tips[tips["tip"] > 5]
print(big_tips)
# Find all female customers who had dinner
dinner_female = tips[(tips['sex'] == 'Female') & (tips['time'] == 'Dinner')]
print(dinner_female.head())
# Using isin to select multiple days
weekend_tips = tips[tips['day'].isin(['Sat', 'Sun'])]
print(weekend_tips.head())
Section 4: GroupBy and Aggregation#
Pandas GroupBy lets us split data into groups and do calculations.
Let us see the average tip by day.
# Average tip by day
avg_tip_by_day = tips.groupby('day')["tip"].mean()
print(avg_tip_by_day)
# Multiple aggregations: average and max tip per sex
grouped = tips.groupby('sex')['tip'].agg(['mean', 'max'])
print(grouped)
# Count the number of tips for each size of group
count_by_size = tips.groupby('size').size()
print(count_by_size)
Section 5: Joining and Merging DataFrames#
Imagine you have more than one table and you need to combine them.
Let us create two small DataFrames and merge them.
# Create two simple DataFrames
left = pd.DataFrame({"id": [1, 2, 3], "name": ["Alice", "Bob", "Cathy"]})
right = pd.DataFrame({"id": [2, 3, 4], "age": [24, 27, 22]})
# Merge on 'id'
merged = pd.merge(left, right, on="id", how="inner")
print(merged)
# Merge Titanic data with new ages DataFrame
ages = pd.DataFrame({"age": [22, 38, 26], "new_col": ["x", "y", "z"]})
merged_titanic = pd.merge(df.head(3), ages, on="age", how="left")
print(merged_titanic)
Section 6: Pivot Tables and Reshaping#
Pivot tables help you reorganize and summarize data quickly.
Let us create a pivot table with the Tips dataset.
# Pivot table: average tip by sex and day
pivot = pd.pivot_table(tips, values='tip', index='sex', columns='day', aggfunc='mean')
print(pivot)
# Melt to go long-form: unwind columns
melted = pd.melt(tips, id_vars=['day'], value_vars=['total_bill', 'tip'])
print(melted.head())
Section 7: Time-Series Handling#
Pandas makes working with dates and times much easier.
Let us create a date column and plot a trend.
import numpy as np
import matplotlib.pyplot as plt
tips['visit_date'] = pd.date_range('2021-01-01', periods=len(tips), freq='D')
# Show what our new column looks like
print(tips[['visit_date', 'total_bill']].head())
# Convert visit_date to datetime if needed
tips['visit_date'] = pd.to_datetime(tips['visit_date'])
plt.figure(figsize=(8,3))
plt.plot(tips['visit_date'], tips['total_bill'], marker='o', linestyle='-')
plt.title('Total Bill Over Time')
plt.xlabel('Visit Date')
plt.ylabel('Total Bill ($)')
plt.tight_layout()
plt.show()
Section 8: Visualization with Pandas#
Let us use Pandas built-in plotting to visualize simple trends.
# Histogram of total bill amounts
tips['total_bill'].plot.hist(bins=20, alpha=0.7)
plt.title('Histogram of Total Bill')
plt.xlabel('Total Bill ($)')
plt.show()
# Box plot of tip by smoker status
tips.boxplot(column='tip', by='smoker')
plt.title('Tip by Smoker')
plt.suptitle('')
plt.xlabel('Smoker')
plt.ylabel('Tip Amount')
plt.show()
Section 9: Mini-Project Part 1 (Exploratory Data Analysis with Titanic)#
Now you will use your skills to look for patterns in the Titanic data.
# Percent of passengers who survived
survival_rate = df['survived'].mean() * 100
print(f"Survival rate: {survival_rate:.1f}%")
# Plot survival rate by sex
df.groupby('sex')['survived'].mean().plot(kind='bar')
plt.title('Survival Rate by Sex')
plt.ylabel('Survival Rate')
plt.ylim(0,1)
plt.show()
# Crosstab for class and survival
cross = pd.crosstab(df['pclass'], df['survived'], normalize='index')
cross.plot(kind='bar', stacked=True)
plt.title('Survival by Ticket Class')
plt.ylabel('Proportion')
plt.xlabel('Ticket Class')
plt.show()
Section 10: Mini-Project Part 2 (Feature Engineering & Deeper Analysis)#
Let us create a new feature to see if families survived together.
# Create a family_size feature
df['family_size'] = df['sibsp'] + df['parch'] + 1
# Check if large families had different survival rates
df['is_large_family'] = df['family_size'] >= 5
print(df.groupby('is_large_family')['survived'].mean())
# Scatter plot of fare vs. age, colored by survival
df.plot.scatter(x='age', y='fare', c='survived', colormap='viridis', alpha=0.7)
plt.title('Fare vs. Age Colored by Survival')
plt.xlabel('Age')
plt.ylabel('Fare')
plt.show()
Section 11: Best Practices and Performance Tips#
Learn how to work efficiently with large data.
# Use .copy() to avoid changing the original DataFrame
tips_copy = tips.copy()
tips_copy['total_bill'] = tips_copy['total_bill'] * 1.1
# Downcast data types for memory savings
tips_small = tips.copy()
tips_small['size'] = pd.to_numeric(tips_small['size'], downcast='unsigned')
print(tips_small['size'].dtype)
Section 12: Troubleshooting Common Pandas Errors#
Let us see some common errors and how to fix them.
# Try to use a column name that does not exist
try:
print(tips['total_tip'])
except KeyError as e:
print("Column not found:", e)
# Check data types before math
try:
result = tips['day'] + 5
except TypeError as e:
print("Cannot add a number to text:", e)
Section 13: Challenge Exercises#
Test your new skills with a fun short challenge!
# EXERCISE: What is the average tip for each time of day?
print(tips.groupby('time')['tip'].mean())
# EXERCISE: Create a DataFrame of only non-smokers at tables of 4 or more
large_nonsmokers = tips[(tips['smoker'] == 'No') & (tips['size'] >= 4)]
print(large_nonsmokers.head())
Recap - What Have We Learned?#
We covered a lot about Pandas today!
- Data loading, exploring, and cleaning
- Conditional selection and grouping
- Joining, reshaping, visualizing, and more
Keep practicing, and these skills will soon feel natural.
See you in the next lesson!
Thank You!#
If you enjoyed this video and want more tutorials, please like and subscribe to our channel.
See you next time!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



