Lesson 3 · Mastering Pandas
Understanding Pandas Series and DataFrames: Essential Python Data Structures for Analysis
In this lesson, we will dive into pandas Series and DataFrames. You will learn how to explore, clean, filter, and analyze real-world data. Let us start our…
- CourseMastering Pandas
- Lesson3 of 44
- Video12 min
- FormatJupyter notebook · 25 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbMastering Pandas: Series and DataFrames#
In this lesson, we will dive into pandas Series and DataFrames.
You will learn how to explore, clean, filter, and analyze real-world data.
Let us start our pandas journey!
import warnings
import numpy as np
np.random.seed(42)
warnings.filterwarnings("ignore") # Suppress warnings for cleaner output
# Data setup (Titanic Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
What is a pandas Series?#
A pandas Series is like one column from a spreadsheet.
It holds data values and an index, which is like row labels.
# Create a Series example
import pandas as pd
ages = pd.Series([22, 38, 26, 35])
print(ages)
# Selecting a column as a Series
survived = df['Survived']
print(type(survived))
print(survived.head())
What is a DataFrame?#
A DataFrame is like an entire spreadsheet, with rows and columns.
Each column is a Series. The DataFrame lets you work with all of them together.
# Getting DataFrame info
print(df.info())
# Basic DataFrame statistics
print(df.describe())
# Viewing DataFrame columns and shape
print("Column names:", df.columns.tolist())
print("Shape:", df.shape)
Cleaning Data: Handling Missing Values#
Often, real-world data is messy.
Let us check for missing values in our Titanic dataset.
# Checking for missing values
print(df.isnull().sum())
# Filling missing Age values with the mean
df['Age'].fillna(df['Age'].mean(), inplace=True)
# Dropping rows with missing Embarked values
df.dropna(subset=['Embarked'], inplace=True)
# Filtering: Passengers under 18
kids = df[df['Age'] < 18]
print(kids[['Name', 'Age', 'Sex']].head())
# Multiple filters: Women first class survivors
women_first_survived = df[(df['Sex'] == 'female') & (df['Pclass'] == 1) & (df['Survived'] == 1)]
print(women_first_survived[['Name', 'Age']].head())
# GroupBy: Survival rates by class
print(df.groupby('Pclass')['Survived'].mean())
# Aggregating statistics for fare by embarkation port
print(df.groupby('Embarked')['Fare'].agg(['mean', 'min', 'max']))
# Value counts for categorical columns
print(df['Sex'].value_counts())
print(df['Embarked'].value_counts())
# Merging: Adding port names to Embarked codes
port_names = {'C': 'Cherbourg', 'Q': 'Queenstown', 'S': 'Southampton'}
df['Port'] = df['Embarked'].map(port_names)
print(df[['Embarked', 'Port']].drop_duplicates())
# Pivot table: Average fare by sex and class
pivot = df.pivot_table('Fare', index='Sex', columns='Pclass', aggfunc='mean')
print(pivot)
# Handling dates: Add a fake 'Date' column for demonstration
import numpy as np
np.random.seed(42)
from datetime import timedelta
base_date = pd.Timestamp('1912-04-01')
df['Date'] = [base_date + timedelta(days=int(x)) for x in np.random.randint(0, 30, len(df))]
print(df[['Name', 'Date']].head())
# Simple time series plot: Number of passengers per day
import matplotlib.pyplot as plt
dcount = df['Date'].value_counts().sort_index()
plt.figure(figsize=(8,3))
plt.plot(dcount.index, dcount.values)
plt.title('Passengers per day')
plt.xlabel('Date')
plt.ylabel('Count')
plt.tight_layout()
plt.show()
# Data visualization: Survival by gender
import seaborn as sns
sns.countplot(data=df, x='Sex', hue='Survived')
plt.title('Survival count by gender')
plt.show()
Mini Project: Feature Engineering#
We will now create a new feature: whether a passenger was traveling alone or not.
# New column: IsAlone
df['IsAlone'] = ((df['SibSp'] == 0) & (df['Parch'] == 0)).astype(int)
print(df[['Name', 'SibSp', 'Parch', 'IsAlone']].head())
# Analyze: Did being alone affect survival?
print(df.groupby('IsAlone')['Survived'].mean())
# Common error: Misspelling column names
try:
df['Aeg']
except KeyError:
print("Column name misspelled! Check spelling.")
# Performance tip: Only load needed columns
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
small_df = pd.read_csv(url, usecols=['Name', 'Age', 'Sex', 'Survived'])
print(small_df.head())
# Let us check what you have learned!
answer = input("Which pandas method previews the top rows of data? ")
if answer.lower() == 'head':
print("Correct! The .head() method previews rows.")
else:
print("Tip: Try df.head() to preview the data!")
Final Recap#
You explored Series and DataFrames. You learned to select, clean, filter, group, and visualize data.
Try applying these pandas skills to your own datasets. Practice makes perfect!
Thank you for learning with us!#
Like this video? Subscribe to our channel for more Python and data science!
Comment below: Which pandas feature do you want to master next?
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



