Mathew K Analytics

Lesson 33 · Python For Time Series

Understanding 7 Key Forecast Accuracy Metrics for Evaluating Models in Python

In this lesson, we will explore how to measure how good a forecast is using Python. We will use real-world time series data, learn several accuracy metrics,…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb
 

Welcome to Forecast Accuracy Metrics in Python#

In this lesson, we will explore how to measure how good a forecast is using Python.

We will use real-world time series data, learn several accuracy metrics, and build up to a simple mini-project.

Whether you want to predict sales or the weather, these tools will help you understand your predictions better!

Let's get started.

# import needed libraries and suppress warnings
import warnings; warnings.filterwarnings("ignore")
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt

What is forecast accuracy?#

Forecast accuracy tells us how close our predictions came to the real values.

If you can measure it, you can improve it!

We will use a real sales dataset for our examples.

# Data setup
url = "https://raw.githubusercontent.com/jbrownlee/Datasets/master/shampoo.csv"
df = pd.read_csv(url)
print("Shape of dataset:", df.shape)
df.head()
Shape of dataset: (36, 2)
Month Sales
0 1-01 266.0
1 1-02 145.9
2 1-03 183.1
3 1-04 119.3
4 1-05 180.3
# Fix column names for easier use
df.columns = ['Month', 'Sales']
df['Sales'] = pd.to_numeric(df['Sales'], errors='coerce')
df = df.dropna()
# Plot sales data
plt.figure(figsize=(8,4))
plt.plot(df['Month'], df['Sales'], marker="o")
plt.title("Monthly Shampoo Sales")
plt.xlabel("Month")
plt.ylabel("Sales")
plt.xticks(rotation=45)
plt.tight_layout()
plt.show()
No description has been provided for this image

Forecasting with a Simple Model#

To measure forecast accuracy, we first need some predictions.

For now, we will predict each month's sales as the previous month's actual value.

This is called a 'naive forecast' and is a common starting point.

# Make naive forecast (shift sales by 1 month)
df['Prediction'] = df['Sales'].shift(1)
print(df[['Month', 'Sales', 'Prediction']].head(10))
  Month  Sales  Prediction
0  1-01  266.0         NaN
1  1-02  145.9       266.0
2  1-03  183.1       145.9
3  1-04  119.3       183.1
4  1-05  180.3       119.3
5  1-06  168.5       180.3
6  1-07  231.8       168.5
7  1-08  224.5       231.8
8  1-09  192.8       224.5
9  1-10  122.9       192.8
# Drop first row (no prediction available)
df = df.dropna()

Introducing Error Metrics#

Forecast error means how far off our predictions were.

We will learn three main metrics today:

  • MAE: Mean Absolute Error
  • RMSE: Root Mean Squared Error
  • MAPE: Mean Absolute Percentage Error
# Calculate MAE
mae = np.mean(np.abs(df['Sales'] - df['Prediction']))
print("Mean Absolute Error:", round(mae, 2))
Mean Absolute Error: 88.22
# Calculate RMSE
rmse = np.sqrt(np.mean((df['Sales'] - df['Prediction']) ** 2))
print("Root Mean Squared Error:", round(rmse, 2))
Root Mean Squared Error: 108.24
# Calculate MAPE (percentage error)
mape = np.mean(np.abs((df['Sales'] - df['Prediction']) / df['Sales'])) * 100
print("Mean Absolute Percentage Error:", round(mape, 2), "%")
Mean Absolute Percentage Error: 30.41 %
# Visualize errors in time
errors = df['Sales'] - df['Prediction']
plt.figure(figsize=(8,4))
plt.plot(df['Month'], errors, marker='o', color='red')
plt.title("Prediction Errors Over Time")
plt.xlabel("Month")
plt.ylabel("Error")
plt.axhline(0, color='black', linewidth=1, linestyle='--')
plt.xticks(rotation=45)
plt.tight_layout()
plt.show()
No description has been provided for this image

Measuring Forecasts: Quick Recap#

  • MAE: Average size of our mistakes.
  • RMSE: Bigger mistakes matter more.
  • MAPE: Mistakes as a percent.

Why do you think different jobs might care about one more than another?

# Try it: Predict next month's sales
guess = input("What is your guess for next month's shampoo sales? ")
real = float(input("The real sales were 330. How close did you get? Type your guess again: "))
print("Your error was:", abs(330 - real))
Your error was: 10.0
 
# More practice: Enter your own prediction and values
pred = float(input("Type any sales prediction number: "))
actual = float(input("Now type the real sales number: "))
print("Absolute error:", abs(actual - pred))
if actual != 0:
    print("Percentage error:", round(abs(actual - pred) / actual * 100, 2), "%")
else:
    print("Percentage error: Not defined for zero actual sales.")
    
Absolute error: 20.0
Percentage error: 6.25 %
 
# Comparison: Random guessing vs naive forecast
np.random.seed(42)
random_guesses = np.random.uniform(low=df['Sales'].min(), high=df['Sales'].max(), size=len(df))
mae_random = np.mean(np.abs(df['Sales'] - random_guesses))
mae_naive = np.mean(np.abs(df['Sales'] - df['Prediction']))
print("Random guess MAE:", round(mae_random, 2))
print("Naive forecast MAE:", round(mae_naive, 2))
Random guess MAE: 166.31
Naive forecast MAE: 88.22

Using Built-in Scoring Functions#

Libraries can calculate errors for you.

We will use scikit-learn's tools. They work the same way, but save typing.

# Use sklearn metrics
from sklearn.metrics import mean_absolute_error, mean_squared_error
mae_sk = mean_absolute_error(df['Sales'], df['Prediction'])
rmse_sk = np.sqrt(mean_squared_error(df['Sales'], df['Prediction']))
print("MAE (sklearn):", round(mae_sk, 2))
print("RMSE (sklearn):", round(rmse_sk, 2))
MAE (sklearn): 88.22
RMSE (sklearn): 108.24
# Mini-project: Compare two forecast methods
df['MeanForecast'] = df['Sales'].rolling(2).mean().shift(1)
mae_naive = np.mean(np.abs(df['Sales'] - df['Prediction']))
mae_mean = np.mean(np.abs(df['Sales'] - df['MeanForecast']))
print("Naive MAE:", round(mae_naive,2))
print("Mean Forecast MAE:", round(mae_mean,2))
Naive MAE: 88.22
Mean Forecast MAE: 64.02
# See forecasts versus actuals
plt.figure(figsize=(9,5))
plt.plot(df['Month'], df['Sales'], label='Actual Sales', marker='o')
plt.plot(df['Month'], df['Prediction'], label='Naive', linestyle='--')
plt.plot(df['Month'], df['MeanForecast'], label='Mean Forecast', linestyle=':')
plt.ylabel('Sales')
plt.xlabel('Month')
plt.title('Actual and Predicted Sales')
plt.legend()
plt.xticks(rotation=45)
plt.tight_layout()
plt.show()
No description has been provided for this image
# Troubleshooting: Dealing with missing or zero sales
df_missing = df.copy()
df_missing.iloc[5, 1] = np.nan
df_missing.iloc[7, 1] = 0
try:
    mape_missing = np.mean(np.abs((df_missing['Sales'] - df_missing['Prediction']) / df_missing['Sales'])) * 100
except Exception as e:
    print("Error calculating MAPE:", e)
else:
    print("MAPE with missing/zero values:", round(mape_missing,2))
    
MAPE with missing/zero values: inf
# Extra tip: Handling outliers
q_low = df['Sales'].quantile(0.01)
q_high = df['Sales'].quantile(0.99)
outliers = df[(df['Sales'] < q_low) | (df['Sales'] > q_high)]
print("Found outliers at:")
print(outliers[['Month', 'Sales']])
Found outliers at:
   Month  Sales
3   1-04  119.3
32  3-09  682.0

Challenge Time!#

  1. Can you invent another rule for forecasting using more months?
  2. Calculate MAE using just those forecast guesses.
  3. Try the whole lesson with a different dataset and see what changes.

Share your ideas or results in the comments!

Recap: What We Learned#

  • What forecast error means.
  • How to calculate MAE, RMSE, and MAPE.
  • Why real sales data is never perfect.
  • Comparing simple and improved prediction models.
  • Visualizing and troubleshooting errors.

Great job!

Thanks for learning with us!#

Practice on your own, try different datasets, and experiment.

If you liked this lesson, hit 'Like' and 'Subscribe' so you never miss new tutorials.

See you next time!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.