Lesson 18 · Python For Time Series
Handling Missing Data in Time Series: Proven Techniques for Accurate Analysis
Discover how to find and fix missing values in time-based data sets, like sales or weather. We will use real airline passenger numbers to practice. No…
- CoursePython For Time Series
- Lesson18 of 30
- Video14 min
- FormatJupyter notebook · 20 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynb
Welcome: Handling Missing Data in Time Series#
Discover how to find and fix missing values in time-based data sets, like sales or weather. We will use real airline passenger numbers to practice. No experience needed just follow along and have fun!
import warnings
warnings.filterwarnings("ignore")
# Let us load the tools we need
import pandas as pd
import matplotlib.pyplot as plt
# We use pandas for data and matplotlib for charts
What are Missing Values?#
Missing values are blanks or special codes where data is not recorded. We need to spot them and handle them before we can use our data. This helps prevent errors and makes our results more trustworthy.
# Data setup: Load the airline passenger data
url = "https://raw.githubusercontent.com/jbrownlee/Datasets/master/airline-passengers.csv"
df = pd.read_csv(url)
# Let us look at the shape and the first few rows
print("Table shape:", df.shape)
print(df.head())
# Make a line chart of the data
plt.figure(figsize=(10,4))
plt.plot(df["Passengers"])
plt.title("Airline Passengers Over Time")
plt.xlabel("Month Number")
plt.ylabel("Passengers")
plt.show()
Creating Missing Data (for Practice)#
Sometimes, data does not have blanks yet. For practice, we will make some missing spots ourselves. This helps us learn how to spot and handle them safely.
# Add missing values in a safe way
import numpy as np
df.loc[2, "Passengers"] = np.nan
df.loc[6, "Passengers"] = np.nan
df.loc[10, "Passengers"] = np.nan
print(df.head(12))
# Find all missing values
print(df.isna())
# Count missing for each column
print(df.isna().sum())
# Show only rows with missing data
missing_rows = df[df.isna().any(axis=1)]
print(missing_rows)
# We use .any(axis=1) to look at each row
Why Handle Missing Values?#
Most Python commands will crash or give wrong results if blanks are there. Next, we will try some ways to fill or skip missing numbers. This makes your time series ready for real modeling or prediction.
# Option 1: Drop rows with missing values
dropped = df.dropna()
print("After drop:")
print(dropped.head(12))
# Option 2: Fill missing with a chosen value
filled = df.fillna(0)
print("After fill:")
print(filled.head(12))
# Option 3: Fill with the average (mean)
mean_value = df["Passengers"].mean()
avg_filled = df.fillna({"Passengers": mean_value})
print("After mean fill:")
print(avg_filled.head(12))
# Interpolate to fill smoothly
interpolated = df.interpolate()
print("After interpolate:")
print(interpolated.head(12))
# Forward fill: carry last value forward
ffilled = df.fillna(method="ffill")
print("Forward filled:")
print(ffilled.head(12))
# Backward fill: fill blanks with the next value
bfilled = df.fillna(method="bfill")
print("Backward filled:")
print(bfilled.head(12))
# What happens if we do nothing? Try using sum.
try:
print(df["Passengers"].sum())
except Exception as e:
print("Error:", e)
# Practice: Enter a value to fill a missing spot
row = int(input("Which row to fix? (type a number, e.g. 2): "))
value = float(input("Fill with what number? "))
df.at[row, "Passengers"] = value
print(df.loc[row])
# Plot again after fixing missing data
plt.plot(df["Passengers"])
plt.title("After Missing Fixed")
plt.xlabel("Month Number")
plt.ylabel("Passengers")
plt.show()
Mini-Project: Fill and Predict Future Passenger Counts#
Try filling all the missing values, then predict what the next month might be if the trend continues. You can use mean, interpolation, or your own fill. Plot your result and look for patterns.
# Solution: Fill by interpolation and predict next month
df_filled = df.interpolate()
last = df_filled["Passengers"].iloc[-1]
second_last = df_filled["Passengers"].iloc[-2]
# Predict next month by extending the last difference
next_passengers = last + (last - second_last)
print("Predicted next month:", next_passengers)
Best Practices for Missing Data#
- Always check for blanks before you do any modeling.
- Try out different fill techniques and think about what makes sense for your business problem.
- Keep track of where you filled or dropped data, so you can explain your results.
# Troubleshooting: What if my fill causes weird results?
problems = df["Passengers"].isna().sum() > 0
if problems:
print("Some missing values still remain! Double-check your fills.")
else:
print("No missing values detected.")
# Extra: Make a mask column showing missing status
df["IsMissing"] = df["Passengers"].isna()
print(df.head(12))
Challenge: Try it Yourself!#
- Add more missing values to random places.
- Try different filling methods for each one.
- Plot before and after to see which fill gives the most natural chart.
Recap: What You Learned#
- What missing values are and why they matter
- How to spot, count, and fill blanks
- Different fill styles for real-world time series data
Keep practicing and you will master time series cleaning!
Thank You! Watch More Python Lessons#
Like this video and subscribe for more fun Python data projects! You will learn new tools each week. See you next time!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



