Mathew K Analytics

Lesson 31 · Python For Time Series

Using Random Forest and XGBoost for Accurate Time Series Forecasting

In this lesson, we will learn how to use two powerful machine learning models, Random Forest and XGBoost, to forecast time series data. Do not worry if…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb
 

Welcome to Time Series Forecasting with Random Forest and XGBoost#

In this lesson, we will learn how to use two powerful machine learning models, Random Forest and XGBoost, to forecast time series data.

Do not worry if you're new we will start with the basics and build up step by step.

By the end, you will know how to prepare time series data, train models, and make predictions!

# Let's start by turning off warnings for a cleaner experience
import warnings
warnings.filterwarnings("ignore")

Step 1: Data setup#

We will use the Shampoo Sales dataset, a classic time series for retail forecasting.

Lets download and preview the data.

# Download the shampoo dataset and check the contents
import pandas as pd
url = "https://raw.githubusercontent.com/jbrownlee/Datasets/master/shampoo.csv"
df = pd.read_csv(url)
print("Data shape:", df.shape)
df.head()
Data shape: (36, 2)
Month Sales
0 1-01 266.0
1 1-02 145.9
2 1-03 183.1
3 1-04 119.3
4 1-05 180.3
# Make a simple line plot to visualize the sales over time
import matplotlib.pyplot as plt
plt.plot(df["Month"], df["Sales"])
plt.xlabel("Month")
plt.ylabel("Sales")
plt.title("Monthly Shampoo Sales")
plt.xticks(rotation=45)
plt.show()
No description has been provided for this image

Step 2: Understanding time series features#

Most machine learning models like Random Forests or XGBoost need the data in a special format, called tabular.

This means we must turn our series into rows of 'features' to help the model learn the patterns.

Lets see how to do this.

# Create 'lag' features for machine learning
df["Sales_prev1"] = df["Sales"].shift(1)
df["Sales_prev2"] = df["Sales"].shift(2)
df = df.dropna()
df.head()
Month Sales Sales_prev1 Sales_prev2
2 1-03 183.1 145.9 266.0
3 1-04 119.3 183.1 145.9
4 1-05 180.3 119.3 183.1
5 1-06 168.5 180.3 119.3
6 1-07 231.8 168.5 180.3

Step 3: Train-test split#

To fairly check our model, we need separate training and testing sets.

For time series, we must use earlier months for training, and later months for testing.

Lets split our data.

# Split the data into train and test sets
train = df.iloc[:-6]
test = df.iloc[-6:]
print("Train shape:", train.shape)
print("Test shape:", test.shape)
Train shape: (28, 4)
Test shape: (6, 4)

Step 4: Random Forest basics#

Random Forest is an ensemble algorithm it builds a group of decision trees and averages their guesses.

It can capture patterns in numbers, like sales over time.

Lets train our first Random Forest model.

# Fit Random Forest to the train set
from sklearn.ensemble import RandomForestRegressor
X_train = train[["Sales_prev1", "Sales_prev2"]]
y_train = train["Sales"]
rf = RandomForestRegressor(random_state=42)
rf.fit(X_train, y_train)
RandomForestRegressor(random_state=42)
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
# Predict sales for the test set and check results
X_test = test[["Sales_prev1", "Sales_prev2"]]
y_test = test["Sales"]
y_pred_rf = rf.predict(X_test)
print("Random Forest predictions:", y_pred_rf)
print("True Sales:", y_test.values)
Random Forest predictions: [403.53  375.08  431.05  369.778 375.099 375.099]
True Sales: [575.5 407.6 682.  475.3 581.3 646.9]
# Plot the predicted vs. actual sales
plt.figure(figsize=(8,4))
plt.plot(test["Month"], y_test, label="Actual")
plt.plot(test["Month"], y_pred_rf, label="Random Forest Predicted")
plt.xlabel("Month")
plt.ylabel("Sales")
plt.title("Random Forest: Actual vs Predicted Sales")
plt.legend()
plt.xticks(rotation=45)
plt.show()
No description has been provided for this image

Step 5: Introduction to XGBoost#

XGBoost is another tree-based algorithm that is often even more accurate.

It builds trees one after another, each time trying to improve.

Lets install and test XGBoost.

# If running locally, you may need to install xgboost first
# Uncomment and run the next line if needed:
# !pip install xgboost
# Train XGBoost on our sales data
from xgboost import XGBRegressor
xgb = XGBRegressor(random_state=42)
xgb.fit(X_train, y_train)
y_pred_xgb = xgb.predict(X_test)
# Compare XGBoost predictions to actual sales
print("XGBoost predictions:", y_pred_xgb)
print("True Sales:", y_test.values)
XGBoost predictions: [407.2409  322.554   437.3997  322.99426 322.554   322.554  ]
True Sales: [575.5 407.6 682.  475.3 581.3 646.9]
# Plot XGBoost and Random Forest together
plt.figure(figsize=(8,4))
plt.plot(test["Month"], y_test, label="Actual")
plt.plot(test["Month"], y_pred_rf, label="Random Forest")
plt.plot(test["Month"], y_pred_xgb, label="XGBoost")
plt.xlabel("Month")
plt.ylabel("Sales")
plt.title("Comparison: Actual vs Random Forest vs XGBoost")
plt.legend()
plt.xticks(rotation=45)
plt.show()
No description has been provided for this image

Step 6: Measuring model performance#

Lets use a metric called Mean Absolute Error (MAE) to see how far off our predictions are from the real sales.

Lower MAE means better forecasting!

We will calculate MAE for both models.

# Evaluate accuracy with MAE
from sklearn.metrics import mean_absolute_error
mae_rf = mean_absolute_error(y_test, y_pred_rf)
mae_xgb = mean_absolute_error(y_test, y_pred_xgb)
print("Random Forest MAE:", mae_rf)
print("XGBoost MAE:", mae_xgb)
Random Forest MAE: 173.1606666666661
XGBoost MAE: 205.55053100585937

Mini-Project: Real-world forecasting#

Imagine you run a store and need to know the next month's shampoo sales.

Lets use our models to predict the following month, using the latest known sales.

# Predict the next month's sales using the final test lags
last_row = test.iloc[-1]
future_lags = [[last_row["Sales_prev1"], last_row["Sales_prev2"]]]
rf_next = rf.predict(future_lags)[0]
xgb_next = xgb.predict(future_lags)[0]
print("Random Forest predicts next month's sales as:", round(rf_next, 2))
print("XGBoost predicts next month's sales as:", round(xgb_next, 2))
Random Forest predicts next month's sales as: 375.1
XGBoost predicts next month's sales as: 322.55

Best practices for time series features#

You can improve your models by adding more features, like the month number, holidays, or year.

But always make sure you use only past data when forecasting the future!

Avoid using features that 'peek' at the answer.

Troubleshooting tips#

  • Are there missing values? Use dropna() or fillna() to clean them.
  • Are your train and test sets mixed? Use iloc to split by row order.
  • Wont fit? Check that your input feature shapes match.

If the error message is about shapes or column names, double-check your code.

# Try getting a feature from user input
print("Enter a value for sales two months ago:")
lag2 = float(input())
print("Enter a value for sales one month ago:")
lag1 = float(input())
guess = rf.predict([[lag1, lag2]])[0]
print("Random Forest would predict:", round(guess, 2))
Enter a value for sales two months ago:
Enter a value for sales one month ago:
Random Forest would predict: 387.77
 

Quick challenge#

Change the number of lags or add a new column (like month as a number).

Train again and see if results improve!

Experimenting is the key to learning.

Recap and next steps#

Today you used Random Forest and XGBoost for time series forecasting.

You learned data setup, feature creation, training, prediction, and evaluation.

Try these tools on new datasets for more practice.

Keep building forecasting helps in business, weather, science, and more!

Thanks for learning with us!#

If you found this useful, please subscribe and tell a friend about the channel.

Happy coding and happy forecasting!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.