Lesson 16 · Python For Machine Learning
Handling Missing Values and Data Imputation Techniques in Python for Machine Learning
In this lesson, you will discover how to handle missing data in Python using simple tools. Missing values are common in real-world datasets, especially in…
- CoursePython For Machine Learning
- Lesson16 of 16
- Video12 min
- FormatJupyter notebook · 22 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynb
Welcome to Python: Handling Missing Values and Imputation#
In this lesson, you will discover how to handle missing data in Python using simple tools.
Missing values are common in real-world datasets, especially in data science.
By the end, you will know how to find, understand, and fix missing data important skills for any project.
Let us get started!
What are missing values?#
Missing values mean some information is not recorded in the data.
Common symbols for missing data include: NaN (Not a Number), None, or empty cells.
We need to handle missing values before we can analyze or model our data.
# Basic setup
import warnings
warnings.filterwarnings("ignore")
import pandas as pd
import numpy as np
Data setup#
We will use the Titanic dataset for this lesson.
It is a real dataset with missing values.
You will learn to load, explore, and work with it step by step.
# Load Titanic dataset straight from a web address
url = "https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv"
df = pd.read_csv(url)
print("Data shape:", df.shape)
df.head()
# Find out how many missing values are in each column
df.isnull().sum()
# Check the percentage of missing values in each column
(df.isnull().mean() * 100).round(2)
# View rows where at least one value is missing
df[df.isnull().any(axis=1)].head()
# Visualize missing values with a simple bar plot
import matplotlib.pyplot as plt
df.isnull().sum().plot(kind="bar")
plt.title("Missing values per column")
plt.ylabel("Count")
plt.show()
Main ways to handle missing values#
Remove rows or columns with missing data
Fill in missing data (imputation)
Which method to use depends on your data and your goals.
# Remove any rows with missing values
df_no_missing = df.dropna()
print("Original shape:", df.shape)
print("Shape after removing missing rows:", df_no_missing.shape)
# Remove any columns with too many missing values (more than 40%)
threshold = len(df) * 0.6
df_few_missing = df.dropna(axis=1, thresh=threshold)
print("Shape after dropping columns:", df_few_missing.shape)
# Fill missing ages with the average age (mean imputation)
mean_age = df["Age"].mean()
df["Age_filled"] = df["Age"].fillna(mean_age)
df[["Age", "Age_filled"]].head(8)
# Fill missing embarked place with the most common value (mode imputation)
mode_embarked = df["Embarked"].mode()[0]
df["Embarked_filled"] = df["Embarked"].fillna(mode_embarked)
df[["Embarked", "Embarked_filled"]].head(8)
# Forward fill: copy previous value to fill missing values
df["Cabin_ffill"] = df["Cabin"].fillna(method="ffill")
df[["Cabin", "Cabin_ffill"]].head(10)
# Interpolate: estimate missing values between numbers
df["Age_interp"] = df["Age"].interpolate()
df[["Age", "Age_interp"]].head(10)
# Advanced: Use SimpleImputer from scikit-learn for column imputation
from sklearn.impute import SimpleImputer
imputer = SimpleImputer(strategy="median")
df["Age_median"] = imputer.fit_transform(df[["Age"]])
df[["Age", "Age_median"]].head(10)
# Checking that there are no missing ages now
df["Age_median"].isnull().sum()
# Optional: ask the user how to fill a missing value
guess_age = input("Imagine you know nothing about a missing person's age. What would you use to fill it? (Type a number): ")
print("You picked:", guess_age)
# Mini project Part 1: Remove all rows with missing Embarked and fill missing Age with the median
df_clean = df.dropna(subset=["Embarked"]).copy()
median_age = df_clean["Age"].median()
df_clean["Age"] = df_clean["Age"].fillna(median_age)
print("Missing values after cleaning:")
print(df_clean.isnull().sum())
# Mini project Part 2: Compare the survival rate before and after cleaning
orig_survival = df["Survived"].mean()
clean_survival = df_clean["Survived"].mean()
print("Original survival rate:", round(orig_survival, 3))
print("After cleaning survival rate:", round(clean_survival, 3))
Troubleshooting common problems#
If your data does not load, check the URL or file path.
If filling methods throw errors, make sure you typed the correct column names.
For big datasets, imputation can take time be patient!
# Extra: flag which rows were imputed in Age column
df["Was_age_missing"] = df["Age"].isnull()
df["Age_filled_with_mean"] = df["Age"].fillna(mean_age)
df[["Age", "Age_filled_with_mean", "Was_age_missing"]].head(10)
# Challenge: Try filling Age with a constant value using input
custom_age = int(input("Type a number to use for all missing ages (try 30): "))
df["Age_filled_custom"] = df["Age"].fillna(custom_age)
df[["Age", "Age_filled_custom"]].head(10)
Recap#
In this lesson, you learned to:
- Detect missing data
- Remove or fill missing values in several ways
- Use mean, median, mode, and your own ideas for imputation
- Spot the effects of data cleaning
You are now ready to work with real-world datasets with confidence!
Want more beginner Python?#
Like and subscribe to our channel for more practical tutorials!
Comment below: What topic should we cover next?
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



