Mathew K Analytics

Lesson 38 · Python For Time Series

Sequence-to-Sequence Forecasting in Python: Deep Learning for Time Series Prediction

Ready to predict the future? In this lesson, you will learn how to teach Python to forecast time series data using a step-by-step approach. We will use…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb
 

Sequence-to-Sequence Forecasting in Python: Beginner Guide#

Ready to predict the future? In this lesson, you will learn how to teach Python to forecast time series data using a step-by-step approach.

We will use real-world examples like retail sales and weather. By the end, you will build a small forecasting model on your own.

Let us explore the basics and have fun while learning!

What is Sequence-to-Sequence (Seq2Seq) Forecasting?#

Seq2Seq forecasting means using data from the past to predict a whole sequence of future values.

Think of it like guessing the next days' weather based on what you observed this week.

# Let us get Python ready for time series forecasting.
import warnings; warnings.filterwarnings("ignore")
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
# Data setup: let us load a real retail sales dataset for forecasting.
url = "https://raw.githubusercontent.com/jbrownlee/Datasets/master/shampoo.csv"
df = pd.read_csv(url)
print("Shape:", df.shape)
print("Preview:")
print(df.head())
Shape: (36, 2)
Preview:
  Month  Sales
0  1-01  266.0
1  1-02  145.9
2  1-03  183.1
3  1-04  119.3
4  1-05  180.3
# Let us plot the sales numbers to see the patterns.
plt.figure(figsize=(8,4))
plt.plot(df["Sales"])
plt.title("Monthly Shampoo Sales")
plt.xlabel("Month")
plt.ylabel("Sales")
plt.show()
No description has been provided for this image

What is a Time Series?#

A time series is just a list of values measured over time. Examples are daily temperatures, monthly sales, or yearly profits. We can use time series like a diary for numbers.

# Fix missing values and clean up column names if needed.
df = df.rename(columns={"Month":"month","Sales":"sales"})
df["sales"] = pd.to_numeric(df["sales"], errors="coerce")
df = df.dropna().reset_index(drop=True)
print(df.head())
  month  sales
0  1-01  266.0
1  1-02  145.9
2  1-03  183.1
3  1-04  119.3
4  1-05  180.3
# Let us create sequences for Seq2Seq forecasting.
def create_sequences(data, input_size=3, output_size=2):
    X, y = [], []
    for i in range(len(data) - input_size - output_size + 1):
        X.append(data[i:i+input_size])
        y.append(data[i+input_size:i+input_size+output_size])
    return np.array(X), np.array(y)
sales_data = df["sales"].values
X, y = create_sequences(sales_data, input_size=3, output_size=2)
print("Input shape:", X.shape, "Output shape:", y.shape)
Input shape: (32, 3) Output shape: (32, 2)
# Look at some sample input-output sequence pairs.
for i in range(3):
    print("Input:", X[i], "-> Output:", y[i])
    
Input: [266.  145.9 183.1] -> Output: [119.3 180.3]
Input: [145.9 183.1 119.3] -> Output: [180.3 168.5]
Input: [183.1 119.3 180.3] -> Output: [168.5 231.8]
# Let us split the data into train and test sets for fair testing.
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42
)
print("Training samples:", X_train.shape[0])
print("Testing samples:", X_test.shape[0])
Training samples: 25
Testing samples: 7
# Build a super simple sequence-to-sequence model: Linear Regression.
from sklearn.linear_model import LinearRegression
model = LinearRegression()
X_train_flat = X_train.reshape(X_train.shape[0], -1)
y_train_flat = y_train.reshape(y_train.shape[0], -1)
model.fit(X_train_flat, y_train_flat)
LinearRegression()
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
# Make predictions on our test set and see results.
X_test_flat = X_test.reshape(X_test.shape[0], -1)
y_pred = model.predict(X_test_flat)
for i in range(3):
    print("Predicted:", np.round(y_pred[i],1), "| True:", y_test[i])
    
Predicted: [545.4 556.5] | True: [682.  475.3]
Predicted: [242.  283.6] | True: [226.  303.6]
Predicted: [427.2 425.3] | True: [439.3 401.3]
# Plot predictions versus true values for a better picture.
plt.figure(figsize=(8,4))
plt.plot(np.arange(len(y_test)), y_test[:,0], label="True First Month")
plt.plot(np.arange(len(y_test)), y_pred[:,0], label="Predicted First Month")
plt.title("Forecast: True vs Predicted (First Month Ahead)")
plt.legend()
plt.show()
No description has been provided for this image
# Calculate a simple accuracy score for our predictions.
from sklearn.metrics import mean_squared_error
mse = mean_squared_error(y_test, y_pred)
print("Mean Squared Error:", round(mse,2))
Mean Squared Error: 5483.28

Recap: You built a sequence-to-sequence forecaster!#

  • You loaded and cleaned real data.
  • Turned it into input and output sequences.
  • Trained a simple model.
  • Scored how well it predicted.

Let us try a real-world challenge next.

# Practice: Try your own sequence. Enter some test sales values.
user_input = input("Type 3 sales values, separated by commas: ")
vals = [float(x.strip()) for x in user_input.split(",")]
vals_arr = np.array(vals).reshape(1,-1)
pred = model.predict(vals_arr)
print("Your input:", vals)
print("Forecasted next two months:", np.round(pred[0],1))
Your input: [120.0, 130.0, 125.0]
Forecasted next two months: [158.3 136. ]
 
# Mini-project: Forecast next month based on the past year.
past_12 = df["sales"].values[-12:]
X_project = past_12[-3:].reshape(1, -1)
future_pred = model.predict(X_project)
print("Recent three months:", past_12[-3:])
print("Predicted next two months:", np.round(future_pred[0],1))
Recent three months: [475.3 581.3 646.9]
Predicted next two months: [577.8 702.7]
# Challenge: Change to a different input size and re-train.
X2, y2 = create_sequences(sales_data, input_size=6, output_size=2)
X2_train, X2_test, y2_train, y2_test = train_test_split(X2, y2, test_size=0.2, random_state=42)
model2 = LinearRegression()
model2.fit(X2_train.reshape(X2_train.shape[0], -1), y2_train.reshape(y2_train.shape[0], -1))
score2 = model2.score(X2_test.reshape(X2_test.shape[0], -1), y2_test)
print("Model accuracy with input_size=6:", round(score2,2))
Model accuracy with input_size=6: 0.68
# Extra tips: Try a more advanced model if you like challenges.
from sklearn.ensemble import RandomForestRegressor
rf = RandomForestRegressor(random_state=42)
rf.fit(X_train_flat, y_train_flat)
rf_pred = rf.predict(X_test_flat)
rf_mse = mean_squared_error(y_test, rf_pred)
print("Random Forest Mean Squared Error:", round(rf_mse,2))
Random Forest Mean Squared Error: 8720.47
# Trouble? Here is a quick check for data shapes.
print("X_train shape:", X_train.shape)
print("y_train shape:", y_train.shape)
print("X_test shape:", X_test.shape)
print("y_test shape:", y_test.shape)
X_train shape: (25, 3)
y_train shape: (25, 2)
X_test shape: (7, 3)
y_test shape: (7, 2)

You did it! Ready for your challenge?#

  • Change the number of steps to predict (output_size) in the earlier function.
  • Plot both the first and second month's predictions.
  • Try this on another open time series dataset.

Let us see your results in the comments!

Recap & Next Steps#

Today you learned to:

  • Load and clean time series data
  • Create input/output sequences
  • Train and test models in Python
  • Score your forecasts and make improvement

You are now ready to explore deeper forecasting ideas!

Like this video? Subscribe for more Python and AI lessons.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.