Mathew K Analytics

Lesson 10 · Mastering Pandas

Introduction to Data Analysis with the Titanic Dataset Using Pandas

Let us dive into a classic real-world dataset: the Titanic passenger list. You will practice loading, cleaning, exploring, and analyzing this famous data…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Mini Project: Titanic Dataset Analysis#

Let us dive into a classic real-world dataset: the Titanic passenger list.

You will practice loading, cleaning, exploring, and analyzing this famous data using pandas.

By the end, you will feel comfortable with core pandas tasks and ready for your own projects.

# Always suppress warnings to keep output tidy
import warnings; warnings.filterwarnings('ignore')
import numpy as np
np.random.seed(42)
# Data setup (Titanic Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
(891, 12)
   PassengerId  Survived  Pclass  \
0            1         0       3   
1            2         1       1   
2            3         1       3   

                                                Name     Sex   Age  SibSp  \
0                            Braund, Mr. Owen Harris    male  22.0      1   
1  Cumings, Mrs. John Bradley (Florence Briggs Th...  female  38.0      1   
2                             Heikkinen, Miss. Laina  female  26.0      0   

   Parch            Ticket     Fare Cabin Embarked  
0      0         A/5 21171   7.2500   NaN        S  
1      0          PC 17599  71.2833   C85        C  
2      0  STON/O2. 3101282   7.9250   NaN        S  
# List Titanic columns and their data types
print(df.dtypes)
PassengerId      int64
Survived         int64
Pclass           int64
Name            object
Sex             object
Age            float64
SibSp            int64
Parch            int64
Ticket          object
Fare           float64
Cabin           object
Embarked        object
dtype: object

Data at a Glance#

  • Rows are passengers.
  • Columns like 'Survived', 'Pclass', 'Sex', 'Age', and others give information about each one.
  • Our goal: answer questions and find patterns.
# How many passengers survived?
survived_count = df['Survived'].value_counts()
print(survived_count)
Survived
0    549
1    342
Name: count, dtype: int64
# Calculate the survival rate
rate = df['Survived'].mean()
print(f'Survival rate: {rate:.2%}')
Survival rate: 38.38%

Cleaning: Checking for Missing Data#

Often data is not perfect. Let us spot what is missing.

# Count missing values in each column
print(df.isnull().sum())
PassengerId      0
Survived         0
Pclass           0
Name             0
Sex              0
Age            177
SibSp            0
Parch            0
Ticket           0
Fare             0
Cabin          687
Embarked         2
dtype: int64
# Fill missing ages with the median value
median_age = df['Age'].median()
df['Age'].fillna(median_age, inplace=True)
print(df['Age'].isnull().sum())
0
# Drop columns we do not need for a simple analysis
cols_to_drop = ['Cabin', 'Ticket', 'Name']
df.drop(columns=cols_to_drop, inplace=True)
print(df.columns)
Index(['PassengerId', 'Survived', 'Pclass', 'Sex', 'Age', 'SibSp', 'Parch',
       'Fare', 'Embarked'],
      dtype='object')

Data Filtering: Women and Children First?#

Let us explore survival rates for females and for children under 16.

# What percent of women survived?
female_survival = df[df['Sex'] == 'female']['Survived'].mean()
print(f'Female survival rate: {female_survival:.2%}')
Female survival rate: 74.20%
# What percent of children under 16 survived?
children_survival = df[df['Age'] < 16]['Survived'].mean()
print(f'Child survival rate (age < 16): {children_survival:.2%}')
Child survival rate (age < 16): 59.04%
# Comparing survival by passenger class
print(df.groupby('Pclass')['Survived'].mean())
Pclass
1    0.629630
2    0.472826
3    0.242363
Name: Survived, dtype: float64

Aggregation Practice#

Try GroupBy aggregations on your own:

  • What is the average fare per class?
  • Are males or females older on average?

Pause the video, try, and resume for answers!

# Mini visualization: survival rate by sex
import matplotlib.pyplot as plt
df.groupby('Sex')['Survived'].mean().plot(kind='bar')
plt.title('Survival Rate by Sex')
plt.ylabel('Survival Rate')
plt.show()
No description has been provided for this image
# Which family groupings had the most survivors?
df['FamilySize'] = df['SibSp'] + df['Parch'] + 1
print(df.groupby('FamilySize')['Survived'].sum().sort_values(ascending=False).head())
FamilySize
1    163
2     89
3     59
4     21
7      4
Name: Survived, dtype: int64
# Create a new feature: Was the fare high?
df['HighFare'] = df['Fare'] > 100
print(df['HighFare'].value_counts())
HighFare
False    838
True      53
Name: count, dtype: int64

Recap: What Have You Practiced?#

  • Loading, inspecting, and summarizing Titanic data
  • Cleaning: handling missing data and dropping columns
  • Filtering by gender, age, and class
  • Grouping and aggregating for useful summaries
  • Creating new features
  • Mini visualizations

Next: Try the challenge below and let us know what else you want to learn!

# Challenge: Try this yourself!
avg_fare_women_firstclass = df[(df['Sex'] == 'female') & (df['Pclass'] == 1)]['Fare'].mean()
print(f'Average fare paid by women in first class: {avg_fare_women_firstclass:.2f}')
Average fare paid by women in first class: 106.13

Nice Work! Practice Next Steps#

  • Experiment by changing the filters, columns, or aggregation methods
  • Write your own 'mini-reports' answering new questions

Keep learning, and see you in the next lesson!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.