Mathew K Analytics

Lesson 10 · Data Mining

Essential Data Preprocessing Techniques Using Titanic and Iris Datasets in Python

Welcome to hands-on data mining! This week focuses on cleaning and exploring the famous Titanic and Iris datasets. You will learn how to get data ready for…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Week 3: Hands-on Data Preprocessing with the Titanic and Iris Datasets#

Welcome to hands-on data mining! This week focuses on cleaning and exploring the famous Titanic and Iris datasets. You will learn how to get data ready for analysis, handle missing values, and start uncovering interesting patterns. Let us dive in and see how we can turn real-world data into reliable insights.

# Suppress Warnings For a Clean Output
import warnings
import numpy as np
np.random.seed(42)
warnings.filterwarnings("ignore")
# Data setup (Titanic Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
titanic = pd.read_csv(url)
print(titanic.shape)
print(titanic.head(3))
(891, 12)
   PassengerId  Survived  Pclass  \
0            1         0       3   
1            2         1       1   
2            3         1       3   

                                                Name     Sex   Age  SibSp  \
0                            Braund, Mr. Owen Harris    male  22.0      1   
1  Cumings, Mrs. John Bradley (Florence Briggs Th...  female  38.0      1   
2                             Heikkinen, Miss. Laina  female  26.0      0   

   Parch            Ticket     Fare Cabin Embarked  
0      0         A/5 21171   7.2500   NaN        S  
1      0          PC 17599  71.2833   C85        C  
2      0  STON/O2. 3101282   7.9250   NaN        S  
# Data setup (Iris Dataset)
url = 'https://raw.githubusercontent.com/uiuc-cse/data-fa14/gh-pages/data/iris.csv'
iris = pd.read_csv(url)
print(iris.shape)
print(iris.head(3))
(150, 5)
   sepal_length  sepal_width  petal_length  petal_width species
0           5.1          3.5           1.4          0.2  setosa
1           4.9          3.0           1.4          0.2  setosa
2           4.7          3.2           1.3          0.2  setosa

Preprocessing: Checking for missing values#

Before analysis, we must check and handle missing data. Missing values can cause problems for machine learning models and skew analysis results.

# Look for missing values in Titanic
titanic.isnull().sum()
PassengerId      0
Survived         0
Pclass           0
Name             0
Sex              0
Age            177
SibSp            0
Parch            0
Ticket           0
Fare             0
Cabin          687
Embarked         2
dtype: int64
# Look for missing values in Iris
iris.isnull().sum()
sepal_length    0
sepal_width     0
petal_length    0
petal_width     0
species         0
dtype: int64

Preprocessing: Filling missing values#

Let us fill in missing values so that our models work smoothly.

# Fill missing ages in Titanic with median age
titanic['Age'] = titanic['Age'].fillna(titanic['Age'].median())

# Fill missing embarked values with most common port
titanic['Embarked'] = titanic['Embarked'].fillna(titanic['Embarked'].mode()[0])
# Confirm there are no more missing values
titanic.isnull().sum()
PassengerId      0
Survived         0
Pclass           0
Name             0
Sex              0
Age              0
SibSp            0
Parch            0
Ticket           0
Fare             0
Cabin          687
Embarked         0
dtype: int64

Exploratory Data Analysis (EDA): Summarizing data#

Let us learn to describe datasets and find interesting patterns.

# Quick summary statistics for Titanic
titanic.describe()
PassengerId Survived Pclass Age SibSp Parch Fare
count 891.000000 891.000000 891.000000 891.000000 891.000000 891.000000 891.000000
mean 446.000000 0.383838 2.308642 29.361582 0.523008 0.381594 32.204208
std 257.353842 0.486592 0.836071 13.019697 1.102743 0.806057 49.693429
min 1.000000 0.000000 1.000000 0.420000 0.000000 0.000000 0.000000
25% 223.500000 0.000000 2.000000 22.000000 0.000000 0.000000 7.910400
50% 446.000000 0.000000 3.000000 28.000000 0.000000 0.000000 14.454200
75% 668.500000 1.000000 3.000000 35.000000 1.000000 0.000000 31.000000
max 891.000000 1.000000 3.000000 80.000000 8.000000 6.000000 512.329200
# Count flower species in Iris dataset
iris['species'].value_counts()
species
setosa        50
versicolor    50
virginica     50
Name: count, dtype: int64
# Plot survival rate by sex in Titanic
import seaborn as sns
import matplotlib.pyplot as plt
sns.barplot(x='Sex', y='Survived', data=titanic)
plt.title('Survival Rate by Sex')
plt.show()
No description has been provided for this image

Data Transformation: Converting Text to Numbers#

Models need numbers, not text. Let us encode the Sex column with numbers instead of words.

# Encode Sex as 0 for male, 1 for female
titanic['Sex'] = titanic['Sex'].map({'male': 0, 'female': 1})
titanic[['Sex']].head(3)
Sex
0 0
1 1
2 1
# One-hot encode the Embarked column
titanic = pd.get_dummies(titanic, columns=['Embarked'], drop_first=True)
titanic.head(3)
PassengerId Survived Pclass Name Sex Age SibSp Parch Ticket Fare Cabin Embarked_Q Embarked_S
0 1 0 3 Braund, Mr. Owen Harris 0 22.0 1 0 A/5 21171 7.2500 NaN False True
1 2 1 1 Cumings, Mrs. John Bradley (Florence Briggs Th... 1 38.0 1 0 PC 17599 71.2833 C85 False False
2 3 1 3 Heikkinen, Miss. Laina 1 26.0 0 0 STON/O2. 3101282 7.9250 NaN False True

Practice: Quick Feature Engineering#

It is often useful to create simple features based on existing columns. Try making a new feature to flag children.

# Create a child flag: 1 if Age < 13, else 0
titanic['Child'] = (titanic['Age'] < 13).astype(int)
titanic[['Age', 'Child']].head(5)
Age Child
0 22.0 0
1 38.0 0
2 26.0 0
3 35.0 0
4 35.0 0
# Practice interactive input: Guess a survivor
name = input("Type the name of a passenger from the preview above (e.g., Allen, Miss. Elisabeth Walton): ")
survived = titanic.loc[titanic['Name'].str.contains(name), 'Survived']
if not survived.empty:
    print("Survived:" if survived.values[0] == 1 else "Did not survive.")
else:
    print("Name not found.")
    
Survived:

Recap: What did we learn?#

This week you practiced fetching, cleaning, transforming, and quickly exploring real-world datasets. You are now prepared to preprocess data for machine learning tasks. Keep experimenting and try your own custom features!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.