Mathew K Analytics

Lesson 30 · Python For Time Series

Linear Regression and Regularization Techniques for Accurate Forecasting in Machine Learning

Welcome! In this lesson, we will explore linear regression and regularization by building simple forecasting models with Python. We will use real-world time…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb
 

Linear Regression and Regularization for Forecasting#

Welcome! In this lesson, we will explore linear regression and regularization by building simple forecasting models with Python. We will use real-world time series data to make predictions and understand how to avoid common pitfalls like overfitting.

Let us get started step by step, even if you are brand new to programming!

# Setup: Turn off unnecessary warnings to keep things tidy
import warnings
warnings.filterwarnings("ignore")

What is Linear Regression?#

Linear regression is a way to find a straight-line relationship between numbers. It helps us make predictions based on known data.

Think of it like drawing the best straight line through some points to predict what comes next!

# Data setup: Let us get some real sales data
import pandas as pd

url = "https://raw.githubusercontent.com/jbrownlee/Datasets/master/shampoo.csv"
data = pd.read_csv(url)

print("Data shape:", data.shape)
print(data.head())
Data shape: (36, 2)
  Month  Sales
0  1-01  266.0
1  1-02  145.9
2  1-03  183.1
3  1-04  119.3
4  1-05  180.3
# Let us plot the sales data to visualize patterns
import matplotlib.pyplot as plt

plt.figure(figsize=(8,4))
plt.plot(data["Sales"], label="Monthly Sales")
plt.title("Shampoo Sales Over Time")
plt.xlabel("Month Number")
plt.ylabel("Sales Volume")
plt.legend()
plt.show()
No description has been provided for this image

Why Do We Need Regularization?#

Regularization is a way to help a model avoid following noise or random bumps in our data.

It is like telling the model to keep things simple, so predictions make sense even with new data!

# Prepare the data for forecasting: use numbers instead of dates
import numpy as np

X = np.arange(len(data)).reshape(-1, 1)
y = data["Sales"].values

print("X shape:", X.shape)
print("y shape:", y.shape)
X shape: (36, 1)
y shape: (36,)
# Split our data into training and testing sets
from sklearn.model_selection import train_test_split

X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.25, random_state=42
)

print("Training size:", X_train.shape[0])
print("Testing size:", X_test.shape[0])
Training size: 27
Testing size: 9
# Build and train a basic linear regression model
from sklearn.linear_model import LinearRegression

lr = LinearRegression()
lr.fit(X_train, y_train)

print("Model learned.")
Model learned.
# Use the model to predict sales for test months
y_pred = lr.predict(X_test)

print("First five real sales:", y_test[:5])
print("First five predicted sales:", y_pred[:5])
First five real sales: [646.9 149.5 315.9 575.5 191.4]
First five predicted sales: [511.10062879 266.29673628 410.95358185 455.46338049 299.67908526]
# Plot the actual vs predicted sales
plt.figure(figsize=(8,4))
plt.scatter(X_test, y_test, color="blue", label="Real Sales")
plt.scatter(X_test, y_pred, color="red", label="Predicted Sales")
plt.xlabel("Month Number")
plt.ylabel("Sales Volume")
plt.title("Real vs Predicted Sales")
plt.legend()
plt.show()
No description has been provided for this image
# Check how well our model did: mean squared error
from sklearn.metrics import mean_squared_error

mse = mean_squared_error(y_test, y_pred)
print(f"Mean Squared Error: {mse:.2f}")
Mean Squared Error: 8794.31

Introducing Regularization#

Regularization keeps our model from following every little bump in the data and avoids overfitting.

Ridge regression and Lasso are two popular regularization methods.

Let us see both in action!

# Use Ridge regression (regularized)
from sklearn.linear_model import Ridge

ridge = Ridge(alpha=1.0)
ridge.fit(X_train, y_train)
y_ridge = ridge.predict(X_test)
mse_ridge = mean_squared_error(y_test, y_ridge)

print(f"Ridge MSE: {mse_ridge:.2f}")
Ridge MSE: 8797.24
# Try Lasso regression for comparison
from sklearn.linear_model import Lasso

lasso = Lasso(alpha=0.5)
lasso.fit(X_train, y_train)
y_lasso = lasso.predict(X_test)
mse_lasso = mean_squared_error(y_test, y_lasso)

print(f"Lasso MSE: {mse_lasso:.2f}")
Lasso MSE: 8797.87
# Ask the user to change alpha values themselves
user_alpha = float(input("Type a new alpha value for Ridge regularization: "))
ridge2 = Ridge(alpha=user_alpha)
ridge2.fit(X_train, y_train)
y_ridge2 = ridge2.predict(X_test)
mse_ridge2 = mean_squared_error(y_test, y_ridge2)
print(f"With alpha = {user_alpha}, Ridge MSE: {mse_ridge2:.2f}")
With alpha = 1.0, Ridge MSE: 8797.24
 
# Mini-project: Forecast next month's sales using the best model
last_month = X.max() + 1
best_model = ridge if mse_ridge < mse and mse_ridge < mse_lasso else lr
next_pred = best_model.predict([[last_month]])
print(f"Predicted sales for next month: {next_pred[0]:.2f}")
Predicted sales for next month: 522.23
# Practice: Ask the user to input a future month number
month_input = int(input("Type which future month you want to predict (e.g., 38): "))
future_pred = best_model.predict([[month_input]])
print(f"Predicted sales for month {month_input}: {future_pred[0]:.2f}")
Predicted sales for month 40: 566.74
 

Best Practices and Common Pitfalls#

  • Always check your model on data it has not seen before.
  • Try several regularization values, not just one.
  • Avoid using too many features unless you have lots of data.
  • Watch out for data errors or missing values.

Simple checks keep your forecasts trustworthy!

# If you get an error: Check for missing or weird data
print("Any missing values?", data.isnull().any().any())
print("Minimum sales:", data["Sales"].min())
print("Maximum sales:", data["Sales"].max())
Any missing values? False
Minimum sales: 119.3
Maximum sales: 682.0
# Extra: Try using more than one feature (advanced)
# For now, we just use month number, but you could add weather, holidays, or promotions if you had them.
print("You can try extra features for fun in your own projects!")
You can try extra features for fun in your own projects!

Challenge Exercise#

Try retraining your Ridge or Lasso model with a different split, or try adding an extra feature if you feel bold.

Then, use your model to forecast three steps ahead!

Can you improve your mean squared error?

Recap: What Did We Learn?#

  • How to load and plot real time series data
  • What linear regression is and how to use it
  • Why regularization is helpful for predictions
  • How to compare models and forecast the future

Every little step makes you more skillful!

Thank You for Learning With Us!#

If you enjoyed this lesson, please give the YouTube video a like, share your comments, and subscribe for more friendly Python lessons.

Keep coding and forecasting!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.