Lesson 29 · Python For Time Series
Supervised Learning for Time Series in Python: Data Preparation and Model Setup
Welcome to your beginner lesson! Today, you will learn how to prepare time series data for supervised learning tasks using Python. You will explore a real…
- CoursePython For Time Series
- Lesson29 of 30
- Video11 min
- FormatJupyter notebook · 18 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynb
Supervised Learning Setup for Time Series in Python#
Welcome to your beginner lesson! Today, you will learn how to prepare time series data for supervised learning tasks using Python.
You will explore a real dataset, create features and labels, and finish with your own mini time series prediction project.
No experience needed let us dive in together!
import warnings; warnings.filterwarnings('ignore')
# Basic imports
import pandas as pd
import matplotlib.pyplot as plt
# These tools help you work with data and charts
What is a Time Series?#
A time series is a sequence of data points ordered over time.
Think of the daily temperature, monthly sales, or number of sunspots each year.
Time order is very important in these problems.
# Data setup
url = 'https://raw.githubusercontent.com/jbrownlee/Datasets/master/shampoo.csv'
df = pd.read_csv(url)
print('Shape:', df.shape)
print(df.head())
# Let us look at a simple plot
plt.figure(figsize=(8,4))
plt.plot(df['Sales'])
plt.xlabel('Month')
plt.ylabel('Sales')
plt.title('Monthly Shampoo Sales')
plt.show()
What is Supervised Learning?#
Supervised learning is where you give the computer past examples with answers.
For time series, it means teaching Python to predict future values from earlier ones.
You must tell Python what columns are "input" and which column is "output".
# Checking if data has missing values
print(df.isnull().sum())
# Rename columns for easy access
df.columns = ['Month', 'Sales']
print(df.columns)
# Convert sales to numbers (sometimes CSV is not read as float)
df['Sales'] = pd.to_numeric(df['Sales'], errors='coerce')
print(df.dtypes)
print('Any missing?', df['Sales'].isnull().any())
Creating Input and Output Columns#
To predict the next value, you make columns for earlier months (inputs) and the "next" month (output).
This process is called making "lags".
Example: to predict sales in May, use sales from April.
# Add a column for previous month's sales (lag-1 feature)
df['Prev_Sales'] = df['Sales'].shift(1)
# Drop first row with missing lag value
df = df.dropna()
df.head()
# Splitting data for training and testing
from sklearn.model_selection import train_test_split
X = df[['Prev_Sales']]
y = df['Sales']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
print('Train samples:', len(X_train))
print('Test samples:', len(X_test))
Fitting a Simple Model#
You can use a linear regression model to discover the relationship between previous month's and this month's sales.
This first model is simple, like drawing a line through the points.
# Fit the model
from sklearn.linear_model import LinearRegression
model = LinearRegression()
model.fit(X_train, y_train)
print('Model trained!')
# Predict sales using the model
y_pred = model.predict(X_test)
print('Predictions:', y_pred[:5])
# Measure how well the model works
from sklearn.metrics import mean_squared_error
mse = mean_squared_error(y_test, y_pred)
print('Mean Squared Error:', mse)
# Plot actual vs predicted values
plt.figure(figsize=(8,4))
plt.plot(y_test.values, label='Actual')
plt.plot(y_pred, label='Predicted')
plt.legend()
plt.title('Actual vs Predicted Sales')
plt.show()
Real-World Time Series Tips#
- Use more 'lags' the past two or three months can help predict better.
- Time order matters do not shuffle time series when training models.
- Real data often has missing rows. Always check before starting.
# Challenge: Try a two-month lag
df['Prev_Sales_2'] = df['Sales'].shift(2)
df2 = df.dropna()
X2 = df2[['Prev_Sales', 'Prev_Sales_2']]
y2 = df2['Sales']
X2_train, X2_test, y2_train, y2_test = train_test_split(X2, y2, test_size=0.2, random_state=42)
model2 = LinearRegression()
model2.fit(X2_train, y2_train)
y2_pred = model2.predict(X2_test)
mse2 = mean_squared_error(y2_test, y2_pred)
print('MSE with two lags:', mse2)
# Mini project: Can you predict future sales based on user input?
input_prev = float(input('What was sales last month? '))
input_prev2 = float(input('And two months ago? '))
user_X = [[input_prev, input_prev2]]
user_pred = model2.predict(user_X)
print('Predicted sales next month:', user_pred[0])
# Troubleshooting: What if your data breaks?
try:
fake_input = float(input('Test: Enter a word or wrong type: '))
except ValueError:
print('Input could not be turned into a number. Please enter digits only.')
Recap#
You learned:
- What time series are and why order matters
- How to load real data and check it
- How to make features using lags
- How to train and test a simple model
- How to predict with your own numbers
- How to handle errors in user input
What Next?#
Try using a different lag, or a bigger dataset. See if you can beat the original error!
If you liked this lesson, please share it and subscribe for more tutorials on Python and data science.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



