Mathew K Analytics

Lesson 18 · Python For Time Series

Handling Missing Data in Time Series: Proven Techniques for Accurate Analysis

Discover how to find and fix missing values in time-based data sets, like sales or weather. We will use real airline passenger numbers to practice. No…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb
 

Welcome: Handling Missing Data in Time Series#

Discover how to find and fix missing values in time-based data sets, like sales or weather. We will use real airline passenger numbers to practice. No experience needed just follow along and have fun!

import warnings
warnings.filterwarnings("ignore")

# Let us load the tools we need
import pandas as pd
import matplotlib.pyplot as plt

# We use pandas for data and matplotlib for charts

What are Missing Values?#

Missing values are blanks or special codes where data is not recorded. We need to spot them and handle them before we can use our data. This helps prevent errors and makes our results more trustworthy.

# Data setup: Load the airline passenger data
url = "https://raw.githubusercontent.com/jbrownlee/Datasets/master/airline-passengers.csv"
df = pd.read_csv(url)

# Let us look at the shape and the first few rows
print("Table shape:", df.shape)
print(df.head())
Table shape: (144, 2)
     Month  Passengers
0  1949-01         112
1  1949-02         118
2  1949-03         132
3  1949-04         129
4  1949-05         121
# Make a line chart of the data
plt.figure(figsize=(10,4))
plt.plot(df["Passengers"])
plt.title("Airline Passengers Over Time")
plt.xlabel("Month Number")
plt.ylabel("Passengers")
plt.show()
No description has been provided for this image

Creating Missing Data (for Practice)#

Sometimes, data does not have blanks yet. For practice, we will make some missing spots ourselves. This helps us learn how to spot and handle them safely.

# Add missing values in a safe way
import numpy as np
df.loc[2, "Passengers"] = np.nan
df.loc[6, "Passengers"] = np.nan
df.loc[10, "Passengers"] = np.nan

print(df.head(12))
      Month  Passengers
0   1949-01       112.0
1   1949-02       118.0
2   1949-03         NaN
3   1949-04       129.0
4   1949-05       121.0
5   1949-06       135.0
6   1949-07         NaN
7   1949-08       148.0
8   1949-09       136.0
9   1949-10       119.0
10  1949-11         NaN
11  1949-12       118.0
# Find all missing values
print(df.isna())

# Count missing for each column
print(df.isna().sum())
     Month  Passengers
0    False       False
1    False       False
2    False        True
3    False       False
4    False       False
..     ...         ...
139  False       False
140  False       False
141  False       False
142  False       False
143  False       False

[144 rows x 2 columns]
Month         0
Passengers    3
dtype: int64
# Show only rows with missing data
missing_rows = df[df.isna().any(axis=1)]
print(missing_rows)

# We use .any(axis=1) to look at each row
      Month  Passengers
2   1949-03         NaN
6   1949-07         NaN
10  1949-11         NaN

Why Handle Missing Values?#

Most Python commands will crash or give wrong results if blanks are there. Next, we will try some ways to fill or skip missing numbers. This makes your time series ready for real modeling or prediction.

# Option 1: Drop rows with missing values
dropped = df.dropna()
print("After drop:")
print(dropped.head(12))
After drop:
      Month  Passengers
0   1949-01       112.0
1   1949-02       118.0
3   1949-04       129.0
4   1949-05       121.0
5   1949-06       135.0
7   1949-08       148.0
8   1949-09       136.0
9   1949-10       119.0
11  1949-12       118.0
12  1950-01       115.0
13  1950-02       126.0
14  1950-03       141.0
# Option 2: Fill missing with a chosen value
filled = df.fillna(0)
print("After fill:")
print(filled.head(12))
After fill:
      Month  Passengers
0   1949-01       112.0
1   1949-02       118.0
2   1949-03         0.0
3   1949-04       129.0
4   1949-05       121.0
5   1949-06       135.0
6   1949-07         0.0
7   1949-08       148.0
8   1949-09       136.0
9   1949-10       119.0
10  1949-11         0.0
11  1949-12       118.0
# Option 3: Fill with the average (mean)
mean_value = df["Passengers"].mean()
avg_filled = df.fillna({"Passengers": mean_value})
print("After mean fill:")
print(avg_filled.head(12))
After mean fill:
      Month  Passengers
0   1949-01  112.000000
1   1949-02  118.000000
2   1949-03  283.539007
3   1949-04  129.000000
4   1949-05  121.000000
5   1949-06  135.000000
6   1949-07  283.539007
7   1949-08  148.000000
8   1949-09  136.000000
9   1949-10  119.000000
10  1949-11  283.539007
11  1949-12  118.000000
# Interpolate to fill smoothly
interpolated = df.interpolate()
print("After interpolate:")
print(interpolated.head(12))
After interpolate:
      Month  Passengers
0   1949-01       112.0
1   1949-02       118.0
2   1949-03       123.5
3   1949-04       129.0
4   1949-05       121.0
5   1949-06       135.0
6   1949-07       141.5
7   1949-08       148.0
8   1949-09       136.0
9   1949-10       119.0
10  1949-11       118.5
11  1949-12       118.0
# Forward fill: carry last value forward
ffilled = df.fillna(method="ffill")
print("Forward filled:")
print(ffilled.head(12))
Forward filled:
      Month  Passengers
0   1949-01       112.0
1   1949-02       118.0
2   1949-03       118.0
3   1949-04       129.0
4   1949-05       121.0
5   1949-06       135.0
6   1949-07       135.0
7   1949-08       148.0
8   1949-09       136.0
9   1949-10       119.0
10  1949-11       119.0
11  1949-12       118.0
# Backward fill: fill blanks with the next value
bfilled = df.fillna(method="bfill")
print("Backward filled:")
print(bfilled.head(12))
Backward filled:
      Month  Passengers
0   1949-01       112.0
1   1949-02       118.0
2   1949-03       129.0
3   1949-04       129.0
4   1949-05       121.0
5   1949-06       135.0
6   1949-07       148.0
7   1949-08       148.0
8   1949-09       136.0
9   1949-10       119.0
10  1949-11       118.0
11  1949-12       118.0
# What happens if we do nothing? Try using sum.
try:
    print(df["Passengers"].sum())
except Exception as e:
    print("Error:", e)
    
39979.0
# Practice: Enter a value to fill a missing spot
row = int(input("Which row to fix? (type a number, e.g. 2): "))
value = float(input("Fill with what number? "))
df.at[row, "Passengers"] = value
print(df.loc[row])
Month         1949-03
Passengers      150.0
Name: 2, dtype: object
 
# Plot again after fixing missing data
plt.plot(df["Passengers"])
plt.title("After Missing Fixed")
plt.xlabel("Month Number")
plt.ylabel("Passengers")
plt.show()
No description has been provided for this image

Mini-Project: Fill and Predict Future Passenger Counts#

Try filling all the missing values, then predict what the next month might be if the trend continues. You can use mean, interpolation, or your own fill. Plot your result and look for patterns.

# Solution: Fill by interpolation and predict next month
df_filled = df.interpolate()
last = df_filled["Passengers"].iloc[-1]
second_last = df_filled["Passengers"].iloc[-2]
# Predict next month by extending the last difference
next_passengers = last + (last - second_last)
print("Predicted next month:", next_passengers)
Predicted next month: 474.0

Best Practices for Missing Data#

  1. Always check for blanks before you do any modeling.
  2. Try out different fill techniques and think about what makes sense for your business problem.
  3. Keep track of where you filled or dropped data, so you can explain your results.
# Troubleshooting: What if my fill causes weird results?
problems = df["Passengers"].isna().sum() > 0
if problems:
    print("Some missing values still remain! Double-check your fills.")
else:
    print("No missing values detected.")
    
Some missing values still remain! Double-check your fills.
# Extra: Make a mask column showing missing status
df["IsMissing"] = df["Passengers"].isna()
print(df.head(12))
      Month  Passengers  IsMissing
0   1949-01       112.0      False
1   1949-02       118.0      False
2   1949-03       150.0      False
3   1949-04       129.0      False
4   1949-05       121.0      False
5   1949-06       135.0      False
6   1949-07         NaN       True
7   1949-08       148.0      False
8   1949-09       136.0      False
9   1949-10       119.0      False
10  1949-11         NaN       True
11  1949-12       118.0      False

Challenge: Try it Yourself!#

  1. Add more missing values to random places.
  2. Try different filling methods for each one.
  3. Plot before and after to see which fill gives the most natural chart.

Recap: What You Learned#

  • What missing values are and why they matter
  • How to spot, count, and fill blanks
  • Different fill styles for real-world time series data

Keep practicing and you will master time series cleaning!

Thank You! Watch More Python Lessons#

Like this video and subscribe for more fun Python data projects! You will learn new tools each week. See you next time!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.