Lesson 2 · Mastering Pandas
Installing Pandas and Setting Up Your Python Environment for Data Science Success
Welcome to your journey into intermediate pandas! In this lesson, you will: Install and verify pandas Set up Jupyter for hands-on coding Load real-world…
- CourseMastering Pandas
- Lesson2 of 44
- Video24 min
- FormatJupyter notebook · 19 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbMastering Pandas: Installation and Environment Setup#
Welcome to your journey into intermediate pandas!
In this lesson, you will:
- Install and verify pandas
- Set up Jupyter for hands-on coding
- Load real-world datasets
- Begin to explore the tools for powerful data analysis
# Check Python version and installed packages
import sys
print('Python version:')
print(sys.version)
import pkg_resources
import numpy as np
np.random.seed(42)
installed_packages = sorted([p.project_name for p in pkg_resources.working_set])
print('pandas' in installed_packages)
# Install pandas if missing (and upgrade pip)
import sys
import subprocess
package = 'pandas'
try:
import pandas
print('pandas is already installed')
except ImportError:
print('pandas not found, installing...')
subprocess.check_call([sys.executable, '-m', 'pip', 'install', '--upgrade', 'pip'])
subprocess.check_call([sys.executable, '-m', 'pip', 'install', 'pandas'])
print('pandas installed successfully!')
# Suppress warnings and import pandas
import warnings
warnings.filterwarnings('ignore')
import pandas as pd
print('pandas version:', pd.__version__)
How to launch Jupyter for hands-on data science#
If you are not running this in Jupyter yet:
- Open your terminal or Anaconda prompt
- Run: jupyter notebook
This will launch a web interface in your browser.
You will write and run code in these interactive cells.
# Data setup (Titanic Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
Exploring a DataFrame in pandas#
Now you are holding a real table of survivors and passengers!
Let us explore the Titanic data so you know what you are working with.
- What columns do you see?
- How many rows are there?
- How might this data be useful to predict survival?
# Summarize the DataFrame
print(df.info())
print(df.describe())
# Preview values in a single column
print(df['Sex'].unique())
print(df['Pclass'].value_counts())
# Clean missing values
print(df.isnull().sum())
df = df.dropna(subset=['Age'])
print('Rows after dropping NA in Age:', df.shape[0])
# Create a new column: Age bucket
df['AgeGroup'] = pd.cut(df['Age'], bins=[0, 12, 18, 30, 50, 80], labels=['Child', 'Teen', 'Young Adult', 'Adult', 'Senior'])
print(df[['Age', 'AgeGroup']].head(6))
# Filter: Only adult females
adults = df[(df['Sex'] == 'female') & (df['Age'] >= 18)]
print(adults[['Name', 'Age', 'Sex']].head(3))
# GroupBy: survival by age group
grouped = df.groupby('AgeGroup')['Survived'].mean()
print(grouped)
# Merge: add a family size column from external list
family_sizes = df[['Name', 'SibSp', 'Parch']].copy()
family_sizes['FamilySize'] = family_sizes['SibSp'] + family_sizes['Parch'] + 1
df = pd.merge(df, family_sizes[['Name', 'FamilySize']], on='Name', how='left')
print(df[['Name', 'FamilySize']].head(3))
# Pivot Table: Survival by class and sex
pivot = df.pivot_table(values='Survived', index='Pclass', columns='Sex', aggfunc='mean')
print(pivot)
# Simple time series example using Flights data
import seaborn as sns
flights = sns.load_dataset('flights')
monthly = flights.groupby('month')['passengers'].sum()
monthly.plot(kind='line', title='Total Airline Passengers by Month')
# Visualize Titanic survival by Sex
import matplotlib.pyplot as plt
df['Survived'].groupby(df['Sex']).mean().plot(kind='bar', color=['skyblue', 'salmon'])
plt.ylabel('Proportion Survived')
plt.title('Titanic Survival Rate by Sex')
plt.show()
# Mini-Project: Quick EDA on Titanic Dataset
print('Average fare by survival:', df.groupby('Survived')['Fare'].mean())
print('Median age by class:', df.groupby('Pclass')['Age'].median())
print('Port of Embarkation counts:')
print(df['Embarked'].value_counts(dropna=False))
# Mini-Project: Feature engineering and correlation
df['FarePerPerson'] = df['Fare'] / df['FamilySize']
print(df[['Fare', 'FamilySize', 'FarePerPerson']].head(3))
print('Correlation with survival:')
print(df[['Survived', 'Fare', 'Age', 'FamilySize']].corr())
# Pandas best practices and performance tips
pd.set_option('display.max_columns', 20)
sample = df.sample(n=100, random_state=42)
print(sample.head(3))
# Common pandas errors and how to fix them
try:
print(df['FakeColumn'].head())
except Exception as e:
print('Error:', str(e))
print('Tip: Check your spelling, or use df.columns to see valid names!')
# Challenge: User filters by AgeGroup
group = input('Choose an AgeGroup to inspect (e.g. Child, Teen, Adult): ')
filtered = df[df['AgeGroup'] == group]
print(filtered[['Name', 'Age', 'Survived']].head())
Quick Recap#
You have installed pandas, loaded a real dataset, and practiced data cleaning, transforming, groupby, joining, pivoting, and visualization.
You even explored basic error handling and user interaction.
With these core tools, you are ready to go further in data science!
What next?#
Practice building your own notebooks with different datasets and ask your own questions.
If you liked this, remember to like, subscribe, and leave a comment!
See you in the next pandas adventure!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



