Lesson 1 · Mastering Pandas
Introduction to Pandas: Essential Tools for Data Analysis in Python
Pandas is one of the most popular Python libraries for data analysis, data cleaning, and transformation. In this lesson, you will learn how to explore,…
- CourseMastering Pandas
- Lesson1 of 44
- Video24 min
- FormatJupyter notebook · 18 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbIntroduction to Pandas and Its Role in Data Analysis#
Pandas is one of the most popular Python libraries for data analysis, data cleaning, and transformation. In this lesson, you will learn how to explore, prepare, and analyze real-world data using pandas. We will use the Titanic dataset to practice key skills such as filtering, aggregating, and visualizing data.
import warnings
warnings.filterwarnings("ignore")
# Import pandas
import pandas as pd
import numpy as np
np.random.seed(42)
# Load Titanic data from a GitHub URL
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
df = pd.read_csv(url)
# Print dataset shape
print(df.shape)
# Show first 3 rows
print(df.head(3))
What is a DataFrame?#
A DataFrame is a two-dimensional table used for storing data. It is similar to a spreadsheet in Excel or a SQL table. Each row is an observation, and each column is a field or variable. Pandas makes it easy to manipulate, filter, and analyze these tables.
# Look at column names and types
print(df.dtypes)
# Show more summary info
print(df.info())
# Basic statistics about numeric columns
print(df.describe())
Data Cleaning: Handling Missing Values#
Real-world data is rarely perfect. We often have to deal with missing or incomplete data. In pandas, we can easily check for missing values and decide how to fill or drop them.
# Count missing values per column
print(df.isnull().sum())
# Drop rows where 'Age' is missing
df_clean = df.dropna(subset=['Age'])
print(df_clean.shape)
# Fill missing values in the 'Embarked' column with the mode
mode_embarked = df['Embarked'].mode()[0]
df['Embarked'] = df['Embarked'].fillna(mode_embarked)
print(df['Embarked'].isnull().sum())
Selecting and Filtering Data#
Pandas makes it easy to filter rows and select specific columns. You can use boolean expressions (comparisons) to get just the data you want.
# Select all female passengers
females = df[df['Sex'] == 'female']
print(females.head(3))
# Find passengers under 18 years old
kids = df[df['Age'] < 18]
print(kids[['Name', 'Age', 'Sex']].head(3))
# Select only the columns 'Name', 'Age', and 'Fare'
simple = df[['Name', 'Age', 'Fare']]
print(simple.head(3))
Grouping and Aggregating Data#
Pandas lets you group data and calculate summary statistics easily. This is powerful for answering comparison questions, such as survival rate by gender.
# Calculate average fare by passenger class
avg_fare = df.groupby('Pclass')['Fare'].mean()
print(avg_fare)
# Survival rate by gender
survival_rate = df.groupby('Sex')['Survived'].mean()
print(survival_rate)
# Count survival by class and gender at the same time
counts = df.groupby(['Pclass', 'Sex'])['Survived'].sum()
print(counts)
Merging and Joining DataFrames#
Sometimes you need to combine information from multiple tables. Pandas makes it easy to join data, similar to SQL joins in databases.
# Simulate joining: add a column with random loyalty points
import numpy as np
np.random.seed(42)
loyalty_points = np.random.randint(0, 100, size=len(df))
df2 = pd.DataFrame({'PassengerId': df['PassengerId'], 'Points': loyalty_points})
merged = pd.merge(df, df2, on='PassengerId')
print(merged[['Name', 'Points']].head(3))
Pivot Tables and Data Reshaping#
Pivot tables let you reorganize your data for easy comparison and quick summaries. Pandas can turn long tables into wide tables for reporting.
# Create a pivot table for survival rate by port and class
pivot = pd.pivot_table(df, values='Survived', index='Embarked', columns='Pclass', aggfunc='mean')
print(pivot)
Time-Series Example: Plotting Age Distribution by Class#
While the Titanic dataset is not a time series, we can explore trends within the data, such as age distribution. Pandas integrates well with plotting libraries for visual analysis.
import matplotlib.pyplot as plt
df.boxplot(column='Age', by='Pclass', grid=False)
plt.title('Age Distribution by Passenger Class')
plt.suptitle('')
plt.xlabel('Class')
plt.ylabel('Age')
plt.show()
Mini-Project: Simple Titanic Data Exploration#
Let us put everything together with a short exercise. Can you find all female passengers in first class, calculate their average age, and plot Fare versus Age?
# Find first class females
first_female = df[(df['Sex'] == 'female') & (df['Pclass'] == 1)]
print(first_female[['Name', 'Age', 'Fare']].head(3))
# Average age
avg_age = first_female['Age'].mean()
print(f"Average age: {avg_age:.2f}")
# Plot Fare vs Age
plt.scatter(first_female['Age'], first_female['Fare'])
plt.xlabel('Age')
plt.ylabel('Fare')
plt.title('Fare vs Age for First Class Females')
plt.show()
Best Practices and Performance Tips#
- Always scan the top and bottom of your data with
.head()and.tail(). - Use
df.info()anddf.describe()to explore data health early. - For large datasets, try to process columns instead of rows for faster performance.
- Prefer chaining methods for concise, readable code.
- Save checkpoints with
to_csv()so you can reload without starting over.
Troubleshooting: Common Pandas Errors#
KeyError: This means the column name you wrote does not exist.ValueError: Sometimes this occurs when shapes of data do not match up.SettingWithCopyWarning: Pandas warns you when your changes might not update the original data.- To fix, double-check your column names and use
.copy()as needed.
# Intentional error: typo in column name
try:
mean_age = df['Agge'].mean()
except KeyError as e:
print('Error:', e)
# Try it yourself: simple filtering exercise with input
name_part = input("Type a few letters from a passenger's name: ")
result = df[df['Name'].str.contains(name_part, case=False)]
print(result[['Name', 'Age', 'Sex', 'Fare']].head(3))
Recap: What You Learned#
- How to load data and check shape and columns
- Cleaning, selecting, filtering, and grouping data
- Making pivot tables and simple charts
- Handling common errors in pandas
- Practice problems for more confidence!
Thank You for Joining!#
If you enjoyed learning pandas with us, please hit Like and Subscribe! We would love to hear your questions or suggestions for future lessons. Try exploring a new dataset using what you have learned today. See you in the next lesson!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



