Lesson 10 · Data Mining
Essential Data Preprocessing Techniques Using Titanic and Iris Datasets in Python
Welcome to hands-on data mining! This week focuses on cleaning and exploring the famous Titanic and Iris datasets. You will learn how to get data ready for…
- CourseData Mining
- Lesson10 of 31
- Video13 min
- FormatJupyter notebook · 14 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbWeek 3: Hands-on Data Preprocessing with the Titanic and Iris Datasets#
Welcome to hands-on data mining! This week focuses on cleaning and exploring the famous Titanic and Iris datasets. You will learn how to get data ready for analysis, handle missing values, and start uncovering interesting patterns. Let us dive in and see how we can turn real-world data into reliable insights.
# Suppress Warnings For a Clean Output
import warnings
import numpy as np
np.random.seed(42)
warnings.filterwarnings("ignore")
# Data setup (Titanic Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
titanic = pd.read_csv(url)
print(titanic.shape)
print(titanic.head(3))
# Data setup (Iris Dataset)
url = 'https://raw.githubusercontent.com/uiuc-cse/data-fa14/gh-pages/data/iris.csv'
iris = pd.read_csv(url)
print(iris.shape)
print(iris.head(3))
Preprocessing: Checking for missing values#
Before analysis, we must check and handle missing data. Missing values can cause problems for machine learning models and skew analysis results.
# Look for missing values in Titanic
titanic.isnull().sum()
# Look for missing values in Iris
iris.isnull().sum()
Preprocessing: Filling missing values#
Let us fill in missing values so that our models work smoothly.
# Fill missing ages in Titanic with median age
titanic['Age'] = titanic['Age'].fillna(titanic['Age'].median())
# Fill missing embarked values with most common port
titanic['Embarked'] = titanic['Embarked'].fillna(titanic['Embarked'].mode()[0])
# Confirm there are no more missing values
titanic.isnull().sum()
Exploratory Data Analysis (EDA): Summarizing data#
Let us learn to describe datasets and find interesting patterns.
# Quick summary statistics for Titanic
titanic.describe()
# Count flower species in Iris dataset
iris['species'].value_counts()
# Plot survival rate by sex in Titanic
import seaborn as sns
import matplotlib.pyplot as plt
sns.barplot(x='Sex', y='Survived', data=titanic)
plt.title('Survival Rate by Sex')
plt.show()
Data Transformation: Converting Text to Numbers#
Models need numbers, not text. Let us encode the Sex column with numbers instead of words.
# Encode Sex as 0 for male, 1 for female
titanic['Sex'] = titanic['Sex'].map({'male': 0, 'female': 1})
titanic[['Sex']].head(3)
# One-hot encode the Embarked column
titanic = pd.get_dummies(titanic, columns=['Embarked'], drop_first=True)
titanic.head(3)
Practice: Quick Feature Engineering#
It is often useful to create simple features based on existing columns. Try making a new feature to flag children.
# Create a child flag: 1 if Age < 13, else 0
titanic['Child'] = (titanic['Age'] < 13).astype(int)
titanic[['Age', 'Child']].head(5)
# Practice interactive input: Guess a survivor
name = input("Type the name of a passenger from the preview above (e.g., Allen, Miss. Elisabeth Walton): ")
survived = titanic.loc[titanic['Name'].str.contains(name), 'Survived']
if not survived.empty:
print("Survived:" if survived.values[0] == 1 else "Did not survive.")
else:
print("Name not found.")
Recap: What did we learn?#
This week you practiced fetching, cleaning, transforming, and quickly exploring real-world datasets. You are now prepared to preprocess data for machine learning tasks. Keep experimenting and try your own custom features!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



