Mathew K Analytics

Lesson 16 · Python For Machine Learning

Handling Missing Values and Data Imputation Techniques in Python for Machine Learning

In this lesson, you will discover how to handle missing data in Python using simple tools. Missing values are common in real-world datasets, especially in…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb
 

Welcome to Python: Handling Missing Values and Imputation#

In this lesson, you will discover how to handle missing data in Python using simple tools.

Missing values are common in real-world datasets, especially in data science.

By the end, you will know how to find, understand, and fix missing data important skills for any project.

Let us get started!

What are missing values?#

Missing values mean some information is not recorded in the data.

Common symbols for missing data include: NaN (Not a Number), None, or empty cells.

We need to handle missing values before we can analyze or model our data.

# Basic setup
import warnings
warnings.filterwarnings("ignore")

import pandas as pd
import numpy as np

Data setup#

We will use the Titanic dataset for this lesson.

It is a real dataset with missing values.

You will learn to load, explore, and work with it step by step.

# Load Titanic dataset straight from a web address
url = "https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv"
df = pd.read_csv(url)
print("Data shape:", df.shape)
df.head()
Data shape: (891, 12)
PassengerId Survived Pclass Name Sex Age SibSp Parch Ticket Fare Cabin Embarked
0 1 0 3 Braund, Mr. Owen Harris male 22.0 1 0 A/5 21171 7.2500 NaN S
1 2 1 1 Cumings, Mrs. John Bradley (Florence Briggs Th... female 38.0 1 0 PC 17599 71.2833 C85 C
2 3 1 3 Heikkinen, Miss. Laina female 26.0 0 0 STON/O2. 3101282 7.9250 NaN S
3 4 1 1 Futrelle, Mrs. Jacques Heath (Lily May Peel) female 35.0 1 0 113803 53.1000 C123 S
4 5 0 3 Allen, Mr. William Henry male 35.0 0 0 373450 8.0500 NaN S
# Find out how many missing values are in each column
df.isnull().sum()
PassengerId      0
Survived         0
Pclass           0
Name             0
Sex              0
Age            177
SibSp            0
Parch            0
Ticket           0
Fare             0
Cabin          687
Embarked         2
dtype: int64
# Check the percentage of missing values in each column
(df.isnull().mean() * 100).round(2)
PassengerId     0.00
Survived        0.00
Pclass          0.00
Name            0.00
Sex             0.00
Age            19.87
SibSp           0.00
Parch           0.00
Ticket          0.00
Fare            0.00
Cabin          77.10
Embarked        0.22
dtype: float64
# View rows where at least one value is missing
df[df.isnull().any(axis=1)].head()
PassengerId Survived Pclass Name Sex Age SibSp Parch Ticket Fare Cabin Embarked
0 1 0 3 Braund, Mr. Owen Harris male 22.0 1 0 A/5 21171 7.2500 NaN S
2 3 1 3 Heikkinen, Miss. Laina female 26.0 0 0 STON/O2. 3101282 7.9250 NaN S
4 5 0 3 Allen, Mr. William Henry male 35.0 0 0 373450 8.0500 NaN S
5 6 0 3 Moran, Mr. James male NaN 0 0 330877 8.4583 NaN Q
7 8 0 3 Palsson, Master. Gosta Leonard male 2.0 3 1 349909 21.0750 NaN S
# Visualize missing values with a simple bar plot
import matplotlib.pyplot as plt
df.isnull().sum().plot(kind="bar")
plt.title("Missing values per column")
plt.ylabel("Count")
plt.show()
No description has been provided for this image

Main ways to handle missing values#

  1. Remove rows or columns with missing data

  2. Fill in missing data (imputation)

Which method to use depends on your data and your goals.

# Remove any rows with missing values
df_no_missing = df.dropna()
print("Original shape:", df.shape)
print("Shape after removing missing rows:", df_no_missing.shape)
Original shape: (891, 12)
Shape after removing missing rows: (183, 12)
# Remove any columns with too many missing values (more than 40%)
threshold = len(df) * 0.6
df_few_missing = df.dropna(axis=1, thresh=threshold)
print("Shape after dropping columns:", df_few_missing.shape)
Shape after dropping columns: (891, 11)
# Fill missing ages with the average age (mean imputation)
mean_age = df["Age"].mean()
df["Age_filled"] = df["Age"].fillna(mean_age)
df[["Age", "Age_filled"]].head(8)
Age Age_filled
0 22.0 22.000000
1 38.0 38.000000
2 26.0 26.000000
3 35.0 35.000000
4 35.0 35.000000
5 NaN 29.699118
6 54.0 54.000000
7 2.0 2.000000
# Fill missing embarked place with the most common value (mode imputation)
mode_embarked = df["Embarked"].mode()[0]
df["Embarked_filled"] = df["Embarked"].fillna(mode_embarked)
df[["Embarked", "Embarked_filled"]].head(8)
Embarked Embarked_filled
0 S S
1 C C
2 S S
3 S S
4 S S
5 Q Q
6 S S
7 S S
# Forward fill: copy previous value to fill missing values
df["Cabin_ffill"] = df["Cabin"].fillna(method="ffill")
df[["Cabin", "Cabin_ffill"]].head(10)
Cabin Cabin_ffill
0 NaN NaN
1 C85 C85
2 NaN C85
3 C123 C123
4 NaN C123
5 NaN C123
6 E46 E46
7 NaN E46
8 NaN E46
9 NaN E46
# Interpolate: estimate missing values between numbers
df["Age_interp"] = df["Age"].interpolate()
df[["Age", "Age_interp"]].head(10)
Age Age_interp
0 22.0 22.0
1 38.0 38.0
2 26.0 26.0
3 35.0 35.0
4 35.0 35.0
5 NaN 44.5
6 54.0 54.0
7 2.0 2.0
8 27.0 27.0
9 14.0 14.0
# Advanced: Use SimpleImputer from scikit-learn for column imputation
from sklearn.impute import SimpleImputer
imputer = SimpleImputer(strategy="median")
df["Age_median"] = imputer.fit_transform(df[["Age"]])
df[["Age", "Age_median"]].head(10)
Age Age_median
0 22.0 22.0
1 38.0 38.0
2 26.0 26.0
3 35.0 35.0
4 35.0 35.0
5 NaN 28.0
6 54.0 54.0
7 2.0 2.0
8 27.0 27.0
9 14.0 14.0
# Checking that there are no missing ages now
df["Age_median"].isnull().sum()
np.int64(0)
# Optional: ask the user how to fill a missing value
guess_age = input("Imagine you know nothing about a missing person's age. What would you use to fill it? (Type a number): ")
print("You picked:", guess_age)
You picked: 28
 
# Mini project Part 1: Remove all rows with missing Embarked and fill missing Age with the median
df_clean = df.dropna(subset=["Embarked"]).copy()
median_age = df_clean["Age"].median()
df_clean["Age"] = df_clean["Age"].fillna(median_age)
print("Missing values after cleaning:")
print(df_clean.isnull().sum())
Missing values after cleaning:
PassengerId          0
Survived             0
Pclass               0
Name                 0
Sex                  0
Age                  0
SibSp                0
Parch                0
Ticket               0
Fare                 0
Cabin              687
Embarked             0
Age_filled           0
Embarked_filled      0
Cabin_ffill          1
Age_interp           0
Age_median           0
dtype: int64
# Mini project Part 2: Compare the survival rate before and after cleaning
orig_survival = df["Survived"].mean()
clean_survival = df_clean["Survived"].mean()
print("Original survival rate:", round(orig_survival, 3))
print("After cleaning survival rate:", round(clean_survival, 3))
Original survival rate: 0.384
After cleaning survival rate: 0.382

Troubleshooting common problems#

If your data does not load, check the URL or file path.

If filling methods throw errors, make sure you typed the correct column names.

For big datasets, imputation can take time be patient!

# Extra: flag which rows were imputed in Age column
df["Was_age_missing"] = df["Age"].isnull()
df["Age_filled_with_mean"] = df["Age"].fillna(mean_age)
df[["Age", "Age_filled_with_mean", "Was_age_missing"]].head(10)
Age Age_filled_with_mean Was_age_missing
0 22.0 22.000000 False
1 38.0 38.000000 False
2 26.0 26.000000 False
3 35.0 35.000000 False
4 35.0 35.000000 False
5 NaN 29.699118 True
6 54.0 54.000000 False
7 2.0 2.000000 False
8 27.0 27.000000 False
9 14.0 14.000000 False
# Challenge: Try filling Age with a constant value using input
custom_age = int(input("Type a number to use for all missing ages (try 30): "))
df["Age_filled_custom"] = df["Age"].fillna(custom_age)
df[["Age", "Age_filled_custom"]].head(10)
Age Age_filled_custom
0 22.0 22.0
1 38.0 38.0
2 26.0 26.0
3 35.0 35.0
4 35.0 35.0
5 NaN 30.0
6 54.0 54.0
7 2.0 2.0
8 27.0 27.0
9 14.0 14.0
 

Recap#

In this lesson, you learned to:

  • Detect missing data
  • Remove or fill missing values in several ways
  • Use mean, median, mode, and your own ideas for imputation
  • Spot the effects of data cleaning

You are now ready to work with real-world datasets with confidence!

Want more beginner Python?#

Like and subscribe to our channel for more practical tutorials!

Comment below: What topic should we cover next?

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.