Lesson 27 · Python For Time Series
Creating Lag and Rolling Features for Time Series Analysis in Python
Welcome! Today we will explore how to create lag and rolling features using Python. These tools help us understand how past values affect the present in…
- CoursePython For Time Series
- Lesson27 of 30
- Video12 min
- FormatJupyter notebook · 19 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynb
Rolling and Lag Features in Time Series with Python#
Welcome! Today we will explore how to create lag and rolling features using Python.
These tools help us understand how past values affect the present in time series data.
Ready? Let's get started!
# Start by ignoring warnings for a smoother experience
import warnings
warnings.filterwarnings("ignore")
What Are Lag and Rolling Features?#
A lag feature helps you look back at previous time steps - like asking, "What was the value last week?"
A rolling feature, also known as moving average, helps you smooth out data by averaging over several past values.
These are useful for forecasting and spotting patterns!
# Data setup: Load the monthly shampoo sales dataset
import pandas as pd
url = "https://raw.githubusercontent.com/jbrownlee/Datasets/master/shampoo.csv"
df = pd.read_csv(url)
df.columns = ["Month", "Sales"]
print("Data shape:", df.shape)
display(df.head())
# Plot the time series to see the sales over time
import matplotlib.pyplot as plt
plt.figure(figsize=(8, 4))
plt.plot(df["Month"], df["Sales"], marker="o")
plt.title("Monthly Shampoo Sales")
plt.xlabel("Month")
plt.ylabel("Sales")
plt.xticks(rotation=45)
plt.tight_layout()
plt.show()
Why Lag Features?#
Imagine you run a store and want to predict next month's sales.
It helps to know what you sold last month and the month before.
Lag features capture this history - making patterns clear for forecasting models.
# Creating a lag feature: sales from the previous month
df["Sales_Lag1"] = df["Sales"].shift(1)
display(df.head())
# Lag features for two and three months ago
df["Sales_Lag2"] = df["Sales"].shift(2)
df["Sales_Lag3"] = df["Sales"].shift(3)
display(df.head())
What Is a Rolling Mean?#
A rolling mean takes an average of sales for a small window, like 3 months at a time.
This helps us see the main trend by removing quick ups and downs.
# Make a rolling average over 3 months
df["Rolling_Mean_3"] = df["Sales"].rolling(window=3).mean()
display(df.head(10))
# Compare sales, lag, and rolling mean on a plot
plt.figure(figsize=(10, 5))
plt.plot(df["Month"], df["Sales"], label="Actual Sales", marker="o")
plt.plot(df["Month"], df["Sales_Lag1"], label="Lag 1 Month", linestyle="--")
plt.plot(df["Month"], df["Rolling_Mean_3"], label="3-Month Avg", linewidth=3, alpha=0.7)
plt.xlabel("Month")
plt.ylabel("Sales")
plt.title("Sales, Lag, and Rolling Average")
plt.legend()
plt.xticks(rotation=45)
plt.tight_layout()
plt.show()
# Identify missing values caused by lagging and rolling
print(df.isnull().sum())
# Drop rows with missing values for modeling
df_clean = df.dropna().reset_index(drop=True)
print(df_clean.head())
Useful Rolling and Lag Variations#
Besides the mean, we can also look at rolling sums, mins, or maxes.
Lag features can use any time gap, not just one month!
# Example: Rolling sum and minimum sales for 3 months
df_clean["Rolling_Sum_3"] = df_clean["Sales"].rolling(window=3).sum()
df_clean["Rolling_Min_3"] = df_clean["Sales"].rolling(window=3).min()
display(df_clean.head())
# Iterating through rows: Print sales and lag for each month
for i, row in df_clean.iterrows():
print(f"Month: {row['Month']}, Sales: {row['Sales']}, Lag-1: {row['Sales_Lag1']}")
# Quick feature creation with a list comprehension for lags 1-3
for l in range(1, 4):
df_clean[f"Sales_Lag{l}_again"] = df_clean["Sales"].shift(l)
display(df_clean.head())
# Filtering: Find months where sales dropped below rolling mean
below_avg = df_clean[df_clean["Sales"] < df_clean["Rolling_Mean_3"]]
display(below_avg[["Month", "Sales", "Rolling_Mean_3"]].head())
# Mini project part 1: Predict the next month's sales using last month's sales (very simple!)
df_clean["Pred_Next_Month"] = df_clean["Sales_Lag1"]
df_clean["Actual_Next_Month"] = df_clean["Sales"].shift(-1)
df_proj = df_clean.dropna().reset_index(drop=True)
print(df_proj[["Sales", "Pred_Next_Month", "Actual_Next_Month"]].head())
# Mini project part 2: Compute mean absolute error (MAE)
mae = (df_proj["Pred_Next_Month"] - df_proj["Actual_Next_Month"]).abs().mean()
print(f"Mean Absolute Error: {mae:.2f}")
# Troubleshooting: What if you see lots of NaN values?
print(df["Sales_Lag1"].isnull().sum(), "missing values in Sales_Lag1")
# Extra tip: Try input() to create your own lag!
n = int(input("How many months back for lag? Enter a number: "))
df[f"Sales_Lag{n}_input"] = df["Sales"].shift(n)
display(df[["Month", "Sales", f"Sales_Lag{n}_input"]].head(n+3))
Challenge Exercise#
Can you create a rolling feature that looks at the last 6 months and grabs the maximum sales?
Hint: Use .rolling(window=6).max()
Try plotting it too!
Recap#
Today you explored lag and rolling features to capture trends and past information.
You used them to make simple predictions and learned how to prepare your data.
Great work!
Thanks and Next Steps#
Want to see more time series tutorials? Subscribe for more lessons and share your favorite time series ideas in the comments!
Happy coding!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



