Lesson 6 · Mastering Pandas
Introduction to Pandas DataFrame: Essential Operations and Attributes Explained
Welcome! In this lesson, you will learn essential pandas DataFrame skills. You will explore, clean, and analyze real data. By the end, you will be confident…
- CourseMastering Pandas
- Lesson6 of 44
- Video28 min
- FormatJupyter notebook · 25 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbMastering Pandas: DataFrame Basics and Operations#
Welcome! In this lesson, you will learn essential pandas DataFrame skills.
You will explore, clean, and analyze real data. By the end, you will be confident using pandas for projects.
Let us begin!
# Suppress warnings for a cleaner notebook
import warnings
import numpy as np
np.random.seed(42)
warnings.filterwarnings("ignore")
Data Setup (Titanic Dataset)#
You will work with the Titanic passenger dataset.
Let us load and preview the data!
# Data setup (Titanic Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
# Learn about the DataFrame's basic attributes
print("Columns:", df.columns.tolist())
print("Indexes:", df.index)
print("Data types:\n", df.dtypes)
# Take a quick statistical look at numeric columns
print(df.describe())
Basic Data Selection: Columns and Rows#
You can select columns and rows in several friendly ways.
Let us try selecting specific data!
# Select a single column as a Series
ages = df['Age']
print(ages.head())
# Select multiple columns as a new DataFrame
df_small = df[['Survived', 'Sex', 'Age']]
print(df_small.head())
# Select rows by index values using loc
print(df.loc[0:4, ['Name', 'Pclass', 'Survived']])
# Select by row number position with iloc
print(df.iloc[0:5, 0:4])
Data Cleaning: Handling Missing Values#
Real data is messy! Pandas makes it easier to handle missing values.
You will quickly check for and fill in missing data.
# Check for missing values in each column
print(df.isnull().sum())
# Fill missing Age values with the median age
age_median = df['Age'].median()
df['Age'].fillna(age_median, inplace=True)
print(df['Age'].isnull().sum())
# Drop rows with missing Embarked values
df = df.dropna(subset=['Embarked'])
print(df.shape)
Filtering Data: Conditional Selection#
It is often useful to focus on certain passengers based on rules.
Let us practice creating smart filters!
# Only look at female passengers
females = df[df['Sex'] == 'female']
print(females.shape)
print(females.head(2))
# Filter: Passengers who paid more than $100
rich = df[df['Fare'] > 100]
print(rich[['Name', 'Fare']].head())
# Multi-condition filter: Male, age < 18, survived
boys = df[(df['Sex'] == 'male') & (df['Age'] < 18) & (df['Survived'] == 1)]
print(boys[['Name', 'Age', 'Survived']])
GroupBy and Aggregation: Summarize by Groups#
Let us see summary statistics for different passenger groups.
Grouping helps you find patterns.
# Average fare by passenger class
print(df.groupby('Pclass')['Fare'].mean())
# Group by sex and survival, then count passengers
print(df.groupby(['Sex', 'Survived'])['PassengerId'].count())
# Aggregating multiple statistics
print(df.groupby('Sex')['Age'].agg(['mean', 'median', 'min', 'max']))
Merging and Joining: Combine DataFrames#
Suppose you have two tables. Pandas makes joining easy.
Let us join a DataFrame of passenger titles to our Titanic dataset.
# Create a simple lookup table of titles
titles = pd.DataFrame({
'Name': df['Name'][:4],
'Title': ['Mr.', 'Mrs.', 'Miss.', 'Master.']
})
print(titles)
# Do a left join to add the Title information
df_joined = df.merge(titles, on='Name', how='left')
print(df_joined[['Name', 'Title']].head(6))
Reshaping: Pivot Tables#
Pivot tables help you summarize and compare groups in a table format.
Let us see how many survived in each class and by sex.
# Make a pivot table for survival by class and sex
table = pd.pivot_table(df, values='PassengerId',
index=['Pclass'],
columns=['Sex'],
aggfunc='count')
print(table)
Visualizing with Pandas: Survival Counts#
A quick bar chart helps you see differences across groups.
Let us visualize survival by passenger class.
# Plot a bar chart: Survival rate by class
import matplotlib.pyplot as plt
df.groupby('Pclass')['Survived'].mean().plot(kind='bar')
plt.ylabel('Survival Rate')
plt.title('Titanic Survival Rate by Passenger Class')
plt.show()
Mini-Project: Exploratory Data Analysis#
Let us try a short challenge. You will investigate age, fare, and survival.
Practice analyzing and plotting!
# Visualize Fare vs. Age by Survival
plt.figure(figsize=(6,4))
colors = ['red' if s == 0 else 'green' for s in df['Survived']]
plt.scatter(df['Age'], df['Fare'], c=colors, alpha=0.5)
plt.xlabel('Age')
plt.ylabel('Fare')
plt.title('Fare vs Age, Colored by Survival')
plt.show()
Best Practices and Performance Tips#
Keep code readable: use comments, clear variable names, and meaningful column names.
Use vectorized pandas methods, not loops, for speed.
Always check your data before and after cleaning.
# Example: Avoid for-loops when possible
df['FareSquared'] = df['Fare'] ** 2
print(df[['Fare', 'FareSquared']].head())
Common Issues and Debugging in Pandas#
Watch for typos in column names pandas will give a KeyError.
If code is slow, check if a loop can be replaced by a vectorized method.
When in doubt, print your DataFrame's shape and head.
# Intentional typo: Try a wrong column name
try:
print(df['Agge'].head())
except KeyError as e:
print("Error:", e)
# Practice: Rename column safely to fix typos
df.rename(columns={'FareSquared': 'Fare_Squared'}, inplace=True)
print(df.columns)
Challenge Yourself!#
Try these:
Find the youngest and oldest passenger in each class.
Create a histogram of fares.
Make a new column flagging passengers as "child" if age < 13.
Pause and try them yourself!
Recap#
You learned to load, explore, clean, join, and visualize data with pandas.
These skills help you work fast and ask deeper questions about real data.
Keep practicing and you will master pandas!
Thank You and Next Steps!#
Thank you for learning with us.
Subscribe for more lessons, and leave your favorite pandas tip in the comments!
Practice, experiment, and you will soon be a pandas pro.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



