Mathew K Analytics

Lesson 3 · Mastering Pandas

Understanding Pandas Series and DataFrames: Essential Python Data Structures for Analysis

In this lesson, we will dive into pandas Series and DataFrames. You will learn how to explore, clean, filter, and analyze real-world data. Let us start our…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Mastering Pandas: Series and DataFrames#

In this lesson, we will dive into pandas Series and DataFrames.

You will learn how to explore, clean, filter, and analyze real-world data.

Let us start our pandas journey!

import warnings
import numpy as np
np.random.seed(42)
warnings.filterwarnings("ignore")  # Suppress warnings for cleaner output
 
 
# Data setup (Titanic Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
(891, 12)
   PassengerId  Survived  Pclass  \
0            1         0       3   
1            2         1       1   
2            3         1       3   

                                                Name     Sex   Age  SibSp  \
0                            Braund, Mr. Owen Harris    male  22.0      1   
1  Cumings, Mrs. John Bradley (Florence Briggs Th...  female  38.0      1   
2                             Heikkinen, Miss. Laina  female  26.0      0   

   Parch            Ticket     Fare Cabin Embarked  
0      0         A/5 21171   7.2500   NaN        S  
1      0          PC 17599  71.2833   C85        C  
2      0  STON/O2. 3101282   7.9250   NaN        S  

What is a pandas Series?#

A pandas Series is like one column from a spreadsheet.

It holds data values and an index, which is like row labels.

# Create a Series example
import pandas as pd
ages = pd.Series([22, 38, 26, 35])
print(ages)
0    22
1    38
2    26
3    35
dtype: int64
# Selecting a column as a Series
survived = df['Survived']
print(type(survived))
print(survived.head())
<class 'pandas.core.series.Series'>
0    0
1    1
2    1
3    1
4    0
Name: Survived, dtype: int64

What is a DataFrame?#

A DataFrame is like an entire spreadsheet, with rows and columns.

Each column is a Series. The DataFrame lets you work with all of them together.

# Getting DataFrame info
print(df.info())
 
 
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 891 entries, 0 to 890
Data columns (total 12 columns):
 #   Column       Non-Null Count  Dtype  
---  ------       --------------  -----  
 0   PassengerId  891 non-null    int64  
 1   Survived     891 non-null    int64  
 2   Pclass       891 non-null    int64  
 3   Name         891 non-null    object 
 4   Sex          891 non-null    object 
 5   Age          714 non-null    float64
 6   SibSp        891 non-null    int64  
 7   Parch        891 non-null    int64  
 8   Ticket       891 non-null    object 
 9   Fare         891 non-null    float64
 10  Cabin        204 non-null    object 
 11  Embarked     889 non-null    object 
dtypes: float64(2), int64(5), object(5)
memory usage: 83.7+ KB
None
# Basic DataFrame statistics
print(df.describe())
       PassengerId    Survived      Pclass         Age       SibSp  \
count   891.000000  891.000000  891.000000  714.000000  891.000000   
mean    446.000000    0.383838    2.308642   29.699118    0.523008   
std     257.353842    0.486592    0.836071   14.526497    1.102743   
min       1.000000    0.000000    1.000000    0.420000    0.000000   
25%     223.500000    0.000000    2.000000   20.125000    0.000000   
50%     446.000000    0.000000    3.000000   28.000000    0.000000   
75%     668.500000    1.000000    3.000000   38.000000    1.000000   
max     891.000000    1.000000    3.000000   80.000000    8.000000   

            Parch        Fare  
count  891.000000  891.000000  
mean     0.381594   32.204208  
std      0.806057   49.693429  
min      0.000000    0.000000  
25%      0.000000    7.910400  
50%      0.000000   14.454200  
75%      0.000000   31.000000  
max      6.000000  512.329200  
# Viewing DataFrame columns and shape
print("Column names:", df.columns.tolist())
print("Shape:", df.shape)
Column names: ['PassengerId', 'Survived', 'Pclass', 'Name', 'Sex', 'Age', 'SibSp', 'Parch', 'Ticket', 'Fare', 'Cabin', 'Embarked']
Shape: (891, 12)

Cleaning Data: Handling Missing Values#

Often, real-world data is messy.

Let us check for missing values in our Titanic dataset.

# Checking for missing values
print(df.isnull().sum())
PassengerId      0
Survived         0
Pclass           0
Name             0
Sex              0
Age            177
SibSp            0
Parch            0
Ticket           0
Fare             0
Cabin          687
Embarked         2
dtype: int64
# Filling missing Age values with the mean
df['Age'].fillna(df['Age'].mean(), inplace=True)
 
 
# Dropping rows with missing Embarked values
df.dropna(subset=['Embarked'], inplace=True)
 
 
# Filtering: Passengers under 18
kids = df[df['Age'] < 18]
print(kids[['Name', 'Age', 'Sex']].head())
                                    Name   Age     Sex
7         Palsson, Master. Gosta Leonard   2.0    male
9    Nasser, Mrs. Nicholas (Adele Achem)  14.0  female
10       Sandstrom, Miss. Marguerite Rut   4.0  female
14  Vestrom, Miss. Hulda Amanda Adolfina  14.0  female
16                  Rice, Master. Eugene   2.0    male
# Multiple filters: Women first class survivors
women_first_survived = df[(df['Sex'] == 'female') & (df['Pclass'] == 1) & (df['Survived'] == 1)]
print(women_first_survived[['Name', 'Age']].head())
                                                 Name        Age
1   Cumings, Mrs. John Bradley (Florence Briggs Th...  38.000000
3        Futrelle, Mrs. Jacques Heath (Lily May Peel)  35.000000
11                           Bonnell, Miss. Elizabeth  58.000000
31     Spencer, Mrs. William Augustus (Marie Eugenie)  29.699118
52           Harper, Mrs. Henry Sleeper (Myna Haxtun)  49.000000
# GroupBy: Survival rates by class
print(df.groupby('Pclass')['Survived'].mean())
Pclass
1    0.626168
2    0.472826
3    0.242363
Name: Survived, dtype: float64
# Aggregating statistics for fare by embarkation port
print(df.groupby('Embarked')['Fare'].agg(['mean', 'min', 'max']))
               mean     min       max
Embarked                             
C         59.954144  4.0125  512.3292
Q         13.276030  6.7500   90.0000
S         27.079812  0.0000  263.0000
# Value counts for categorical columns
print(df['Sex'].value_counts())
print(df['Embarked'].value_counts())
Sex
male      577
female    312
Name: count, dtype: int64
Embarked
S    644
C    168
Q     77
Name: count, dtype: int64
# Merging: Adding port names to Embarked codes
port_names = {'C': 'Cherbourg', 'Q': 'Queenstown', 'S': 'Southampton'}
df['Port'] = df['Embarked'].map(port_names)
print(df[['Embarked', 'Port']].drop_duplicates())
  Embarked         Port
0        S  Southampton
1        C    Cherbourg
5        Q   Queenstown
# Pivot table: Average fare by sex and class
pivot = df.pivot_table('Fare', index='Sex', columns='Pclass', aggfunc='mean')
print(pivot)
Pclass           1          2          3
Sex                                     
female  106.693750  21.970121  16.118810
male     67.226127  19.741782  12.661633
# Handling dates: Add a fake 'Date' column for demonstration
import numpy as np
np.random.seed(42)
from datetime import timedelta
base_date = pd.Timestamp('1912-04-01')
df['Date'] = [base_date + timedelta(days=int(x)) for x in np.random.randint(0, 30, len(df))]
print(df[['Name', 'Date']].head())
                                                Name       Date
0                            Braund, Mr. Owen Harris 1912-04-07
1  Cumings, Mrs. John Bradley (Florence Briggs Th... 1912-04-20
2                             Heikkinen, Miss. Laina 1912-04-29
3       Futrelle, Mrs. Jacques Heath (Lily May Peel) 1912-04-15
4                           Allen, Mr. William Henry 1912-04-11
# Simple time series plot: Number of passengers per day
import matplotlib.pyplot as plt
dcount = df['Date'].value_counts().sort_index()
plt.figure(figsize=(8,3))
plt.plot(dcount.index, dcount.values)
plt.title('Passengers per day')
plt.xlabel('Date')
plt.ylabel('Count')
plt.tight_layout()
plt.show()
No description has been provided for this image
# Data visualization: Survival by gender
import seaborn as sns
sns.countplot(data=df, x='Sex', hue='Survived')
plt.title('Survival count by gender')
plt.show()
No description has been provided for this image

Mini Project: Feature Engineering#

We will now create a new feature: whether a passenger was traveling alone or not.

# New column: IsAlone
df['IsAlone'] = ((df['SibSp'] == 0) & (df['Parch'] == 0)).astype(int)
print(df[['Name', 'SibSp', 'Parch', 'IsAlone']].head())
                                                Name  SibSp  Parch  IsAlone
0                            Braund, Mr. Owen Harris      1      0        0
1  Cumings, Mrs. John Bradley (Florence Briggs Th...      1      0        0
2                             Heikkinen, Miss. Laina      0      0        1
3       Futrelle, Mrs. Jacques Heath (Lily May Peel)      1      0        0
4                           Allen, Mr. William Henry      0      0        1
# Analyze: Did being alone affect survival?
print(df.groupby('IsAlone')['Survived'].mean())
IsAlone
0    0.505650
1    0.300935
Name: Survived, dtype: float64
# Common error: Misspelling column names
try:
    df['Aeg']
except KeyError:
    print("Column name misspelled! Check spelling.")
    
Column name misspelled! Check spelling.
# Performance tip: Only load needed columns
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
small_df = pd.read_csv(url, usecols=['Name', 'Age', 'Sex', 'Survived'])
print(small_df.head())
   Survived                                               Name     Sex   Age
0         0                            Braund, Mr. Owen Harris    male  22.0
1         1  Cumings, Mrs. John Bradley (Florence Briggs Th...  female  38.0
2         1                             Heikkinen, Miss. Laina  female  26.0
3         1       Futrelle, Mrs. Jacques Heath (Lily May Peel)  female  35.0
4         0                           Allen, Mr. William Henry    male  35.0
# Let us check what you have learned!
answer = input("Which pandas method previews the top rows of data? ")
if answer.lower() == 'head':
    print("Correct! The .head() method previews rows.")
else:
    print("Tip: Try df.head() to preview the data!")
    
Correct! The .head() method previews rows.

Final Recap#

You explored Series and DataFrames. You learned to select, clean, filter, group, and visualize data.

Try applying these pandas skills to your own datasets. Practice makes perfect!

Thank you for learning with us!#

Like this video? Subscribe to our channel for more Python and data science!

Comment below: Which pandas feature do you want to master next?

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.