Mathew K Analytics

Lesson 2 · Mastering Pandas

Installing Pandas and Setting Up Your Python Environment for Data Science Success

Welcome to your journey into intermediate pandas! In this lesson, you will: Install and verify pandas Set up Jupyter for hands-on coding Load real-world…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Mastering Pandas: Installation and Environment Setup#

Welcome to your journey into intermediate pandas!

In this lesson, you will:

  • Install and verify pandas
  • Set up Jupyter for hands-on coding
  • Load real-world datasets
  • Begin to explore the tools for powerful data analysis
# Check Python version and installed packages
import sys
print('Python version:')
print(sys.version)
 
import pkg_resources
import numpy as np
np.random.seed(42)
installed_packages = sorted([p.project_name for p in pkg_resources.working_set])
print('pandas' in installed_packages)
Python version:
3.12.1 (tags/v3.12.1:2305ca5, Dec  7 2023, 22:03:25) [MSC v.1937 64 bit (AMD64)]
C:\Users\makmw\AppData\Local\Temp\ipykernel_50756\1572279022.py:6: UserWarning: pkg_resources is deprecated as an API. See https://setuptools.pypa.io/en/latest/pkg_resources.html. The pkg_resources package is slated for removal as early as 2025-11-30. Refrain from using this package or pin to Setuptools<81.
  import pkg_resources
True
# Install pandas if missing (and upgrade pip)
import sys
import subprocess
package = 'pandas'
try:
    import pandas
    print('pandas is already installed')
except ImportError:
    print('pandas not found, installing...')
    subprocess.check_call([sys.executable, '-m', 'pip', 'install', '--upgrade', 'pip'])
    subprocess.check_call([sys.executable, '-m', 'pip', 'install', 'pandas'])
    print('pandas installed successfully!')
    
pandas is already installed
# Suppress warnings and import pandas
import warnings
warnings.filterwarnings('ignore')
import pandas as pd
print('pandas version:', pd.__version__)
pandas version: 2.3.0

How to launch Jupyter for hands-on data science#

If you are not running this in Jupyter yet:

  • Open your terminal or Anaconda prompt
  • Run: jupyter notebook

This will launch a web interface in your browser.

You will write and run code in these interactive cells.

# Data setup (Titanic Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
(891, 12)
   PassengerId  Survived  Pclass  \
0            1         0       3   
1            2         1       1   
2            3         1       3   

                                                Name     Sex   Age  SibSp  \
0                            Braund, Mr. Owen Harris    male  22.0      1   
1  Cumings, Mrs. John Bradley (Florence Briggs Th...  female  38.0      1   
2                             Heikkinen, Miss. Laina  female  26.0      0   

   Parch            Ticket     Fare Cabin Embarked  
0      0         A/5 21171   7.2500   NaN        S  
1      0          PC 17599  71.2833   C85        C  
2      0  STON/O2. 3101282   7.9250   NaN        S  

Exploring a DataFrame in pandas#

Now you are holding a real table of survivors and passengers!

Let us explore the Titanic data so you know what you are working with.

  • What columns do you see?
  • How many rows are there?
  • How might this data be useful to predict survival?
# Summarize the DataFrame
print(df.info())
print(df.describe())
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 891 entries, 0 to 890
Data columns (total 12 columns):
 #   Column       Non-Null Count  Dtype  
---  ------       --------------  -----  
 0   PassengerId  891 non-null    int64  
 1   Survived     891 non-null    int64  
 2   Pclass       891 non-null    int64  
 3   Name         891 non-null    object 
 4   Sex          891 non-null    object 
 5   Age          714 non-null    float64
 6   SibSp        891 non-null    int64  
 7   Parch        891 non-null    int64  
 8   Ticket       891 non-null    object 
 9   Fare         891 non-null    float64
 10  Cabin        204 non-null    object 
 11  Embarked     889 non-null    object 
dtypes: float64(2), int64(5), object(5)
memory usage: 83.7+ KB
None
       PassengerId    Survived      Pclass         Age       SibSp  \
count   891.000000  891.000000  891.000000  714.000000  891.000000   
mean    446.000000    0.383838    2.308642   29.699118    0.523008   
std     257.353842    0.486592    0.836071   14.526497    1.102743   
min       1.000000    0.000000    1.000000    0.420000    0.000000   
25%     223.500000    0.000000    2.000000   20.125000    0.000000   
50%     446.000000    0.000000    3.000000   28.000000    0.000000   
75%     668.500000    1.000000    3.000000   38.000000    1.000000   
max     891.000000    1.000000    3.000000   80.000000    8.000000   

            Parch        Fare  
count  891.000000  891.000000  
mean     0.381594   32.204208  
std      0.806057   49.693429  
min      0.000000    0.000000  
25%      0.000000    7.910400  
50%      0.000000   14.454200  
75%      0.000000   31.000000  
max      6.000000  512.329200  
# Preview values in a single column
print(df['Sex'].unique())
print(df['Pclass'].value_counts())
['male' 'female']
Pclass
3    491
1    216
2    184
Name: count, dtype: int64
# Clean missing values
print(df.isnull().sum())
df = df.dropna(subset=['Age'])
print('Rows after dropping NA in Age:', df.shape[0])
PassengerId      0
Survived         0
Pclass           0
Name             0
Sex              0
Age            177
SibSp            0
Parch            0
Ticket           0
Fare             0
Cabin          687
Embarked         2
dtype: int64
Rows after dropping NA in Age: 714
# Create a new column: Age bucket
df['AgeGroup'] = pd.cut(df['Age'], bins=[0, 12, 18, 30, 50, 80], labels=['Child', 'Teen', 'Young Adult', 'Adult', 'Senior'])
print(df[['Age', 'AgeGroup']].head(6))
    Age     AgeGroup
0  22.0  Young Adult
1  38.0        Adult
2  26.0  Young Adult
3  35.0        Adult
4  35.0        Adult
6  54.0       Senior
# Filter: Only adult females
adults = df[(df['Sex'] == 'female') & (df['Age'] >= 18)]
print(adults[['Name', 'Age', 'Sex']].head(3))
                                                Name   Age     Sex
1  Cumings, Mrs. John Bradley (Florence Briggs Th...  38.0  female
2                             Heikkinen, Miss. Laina  26.0  female
3       Futrelle, Mrs. Jacques Heath (Lily May Peel)  35.0  female
# GroupBy: survival by age group
grouped = df.groupby('AgeGroup')['Survived'].mean()
print(grouped)
AgeGroup
Child          0.579710
Teen           0.428571
Young Adult    0.355556
Adult          0.423237
Senior         0.343750
Name: Survived, dtype: float64
# Merge: add a family size column from external list
family_sizes = df[['Name', 'SibSp', 'Parch']].copy()
family_sizes['FamilySize'] = family_sizes['SibSp'] + family_sizes['Parch'] + 1
df = pd.merge(df, family_sizes[['Name', 'FamilySize']], on='Name', how='left')
print(df[['Name', 'FamilySize']].head(3))
                                                Name  FamilySize
0                            Braund, Mr. Owen Harris           2
1  Cumings, Mrs. John Bradley (Florence Briggs Th...           2
2                             Heikkinen, Miss. Laina           1
# Pivot Table: Survival by class and sex
pivot = df.pivot_table(values='Survived', index='Pclass', columns='Sex', aggfunc='mean')
print(pivot)
Sex       female      male
Pclass                    
1       0.964706  0.396040
2       0.918919  0.151515
3       0.460784  0.150198
# Simple time series example using Flights data
import seaborn as sns
flights = sns.load_dataset('flights')
monthly = flights.groupby('month')['passengers'].sum()
monthly.plot(kind='line', title='Total Airline Passengers by Month')
<Axes: title={'center': 'Total Airline Passengers by Month'}, xlabel='month'>
No description has been provided for this image
# Visualize Titanic survival by Sex
import matplotlib.pyplot as plt
df['Survived'].groupby(df['Sex']).mean().plot(kind='bar', color=['skyblue', 'salmon'])
plt.ylabel('Proportion Survived')
plt.title('Titanic Survival Rate by Sex')
plt.show()
No description has been provided for this image
# Mini-Project: Quick EDA on Titanic Dataset
print('Average fare by survival:', df.groupby('Survived')['Fare'].mean())
print('Median age by class:', df.groupby('Pclass')['Age'].median())
print('Port of Embarkation counts:')
print(df['Embarked'].value_counts(dropna=False))
Average fare by survival: Survived
0    22.965456
1    51.843205
Name: Fare, dtype: float64
Median age by class: Pclass
1    37.0
2    29.0
3    24.0
Name: Age, dtype: float64
Port of Embarkation counts:
Embarked
S      554
C      130
Q       28
NaN      2
Name: count, dtype: int64
# Mini-Project: Feature engineering and correlation
df['FarePerPerson'] = df['Fare'] / df['FamilySize']
print(df[['Fare', 'FamilySize', 'FarePerPerson']].head(3))
print('Correlation with survival:')
print(df[['Survived', 'Fare', 'Age', 'FamilySize']].corr())
      Fare  FamilySize  FarePerPerson
0   7.2500           2        3.62500
1  71.2833           2       35.64165
2   7.9250           1        7.92500
Correlation with survival:
            Survived      Fare       Age  FamilySize
Survived    1.000000  0.268189 -0.077221    0.042787
Fare        0.268189  1.000000  0.096067    0.204640
Age        -0.077221  0.096067  1.000000   -0.301914
FamilySize  0.042787  0.204640 -0.301914    1.000000
# Pandas best practices and performance tips
pd.set_option('display.max_columns', 20)
sample = df.sample(n=100, random_state=42)
print(sample.head(3))
     PassengerId  Survived  Pclass  \
120          150         0       2   
329          408         1       2   
39            54         1       2   

                                                  Name     Sex   Age  SibSp  \
120                  Byles, Rev. Thomas Roussel Davids    male  42.0      0   
329                     Richards, Master. William Rowe    male   3.0      1   
39   Faunthorpe, Mrs. Lizzie (Elizabeth Anne Wilkin...  female  29.0      1   

     Parch  Ticket   Fare Cabin Embarked     AgeGroup  FamilySize  \
120      0  244310  13.00   NaN        S        Adult           1   
329      1   29106  18.75   NaN        S        Child           3   
39       0    2926  26.00   NaN        S  Young Adult           2   

     FarePerPerson  
120          13.00  
329           6.25  
39           13.00  
# Common pandas errors and how to fix them
try:
    print(df['FakeColumn'].head())
except Exception as e:
    print('Error:', str(e))
    print('Tip: Check your spelling, or use df.columns to see valid names!')
    
Error: 'FakeColumn'
Tip: Check your spelling, or use df.columns to see valid names!
# Challenge: User filters by AgeGroup
group = input('Choose an AgeGroup to inspect (e.g. Child, Teen, Adult): ')
filtered = df[df['AgeGroup'] == group]
print(filtered[['Name', 'Age', 'Survived']].head())
                                                 Name   Age  Survived
1   Cumings, Mrs. John Bradley (Florence Briggs Th...  38.0         1
3        Futrelle, Mrs. Jacques Heath (Lily May Peel)  35.0         1
4                            Allen, Mr. William Henry  35.0         0
12                        Andersson, Mr. Anders Johan  39.0         0
16  Vander Planke, Mrs. Julius (Emelia Maria Vande...  31.0         0

Quick Recap#

You have installed pandas, loaded a real dataset, and practiced data cleaning, transforming, groupby, joining, pivoting, and visualization.

You even explored basic error handling and user interaction.

With these core tools, you are ready to go further in data science!

What next?#

Practice building your own notebooks with different datasets and ask your own questions.

If you liked this, remember to like, subscribe, and leave a comment!

See you in the next pandas adventure!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.