Mathew K Analytics

Lesson 33 · Python Fundamentals

Understanding Ridge, Lasso, and ElasticNet Regularization in Python for Regression Analysis

Today, you will learn how Ridge, Lasso, and ElasticNet make your machine learning models better. No experience neededlet us start together and explore the…

⬇ Download notebookOpen in Colab ↗

📓 Full notebook

Download .ipynb
 

Welcome to Python Regularization!#

Today, you will learn how Ridge, Lasso, and ElasticNet make your machine learning models better.

No experience neededlet us start together and explore the basics of regularization step by step!

What is Regularization?#

Regularization helps models avoid 'overfitting'when a model remembers too much from the training data and performs badly on new data.

Ridge and Lasso are two popular regularization methods, and ElasticNet mixes both.

Let us see how it works in Python!

# Import warning filter to keep our notebook clean
import warnings
warnings.filterwarnings("ignore")
# Data setup: let us use the California Housing dataset for regression
from sklearn.datasets import fetch_california_housing
import pandas as pd

data = fetch_california_housing(as_frame=True)
df = data.frame
print("Shape of the data:", df.shape)
df.head()
Shape of the data: (20640, 9)
MedInc HouseAge AveRooms AveBedrms Population AveOccup Latitude Longitude MedHouseVal
0 8.3252 41.0 6.984127 1.023810 322.0 2.555556 37.88 -122.23 4.526
1 8.3014 21.0 6.238137 0.971880 2401.0 2.109842 37.86 -122.22 3.585
2 7.2574 52.0 8.288136 1.073446 496.0 2.802260 37.85 -122.24 3.521
3 5.6431 52.0 5.817352 1.073059 558.0 2.547945 37.85 -122.25 3.413
4 3.8462 52.0 6.281853 1.081081 565.0 2.181467 37.85 -122.25 3.422
# Let us check the target distribution
df["MedHouseVal"].describe()
count    20640.000000
mean         2.068558
std          1.153956
min          0.149990
25%          1.196000
50%          1.797000
75%          2.647250
max          5.000010
Name: MedHouseVal, dtype: float64

Splitting Data: Train and Test Sets#

We want to train our model on one set, and check if it works on new data.

Let us split the data into training (used to learn) and test (used to check).

# Split data: features and target
X = df.drop(columns=["MedHouseVal"])
y = df["MedHouseVal"]

# Split into train and test with 20 percent for testing
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
print("Train shape:", X_train.shape, "Test shape:", X_test.shape)
Train shape: (16512, 8) Test shape: (4128, 8)

Plain Linear Regression (No Regularization Yet)#

First, let us see what happens if we build a simple linear model with no regularization.

This is the starting point before applying Ridge or Lasso.

# Linear regression model
from sklearn.linear_model import LinearRegression
linreg = LinearRegression()
linreg.fit(X_train, y_train)
score_train = linreg.score(X_train, y_train)
score_test = linreg.score(X_test, y_test)
print("Train R^2:", round(score_train, 3), "Test R^2:", round(score_test, 3))
Train R^2: 0.613 Test R^2: 0.576

Ridge Regression (L2 Regularization)#

Ridge helps stop the model from caring too much about any one feature.

It adds a penalty if the model's weights get too big.

# Ridge regression with default strength
from sklearn.linear_model import Ridge
ridge = Ridge(alpha=1.0)
ridge.fit(X_train, y_train)
ridge_train = ridge.score(X_train, y_train)
ridge_test = ridge.score(X_test, y_test)
print("Ridge Train R^2:", round(ridge_train, 3), "Ridge Test R^2:", round(ridge_test, 3))
Ridge Train R^2: 0.613 Ridge Test R^2: 0.576

Lasso Regression (L1 Regularization)#

Lasso is another way to keep models simple.

It can shrink some weights all the way to zero, removing unhelpful features.

# Lasso regression with default strength
from sklearn.linear_model import Lasso
lasso = Lasso(alpha=1.0)
lasso.fit(X_train, y_train)
lasso_train = lasso.score(X_train, y_train)
lasso_test = lasso.score(X_test, y_test)
print("Lasso Train R^2:", round(lasso_train, 3), "Lasso Test R^2:", round(lasso_test, 3))
Lasso Train R^2: 0.29 Lasso Test R^2: 0.284

ElasticNet: Mix of Ridge and Lasso#

ElasticNet combines Ridge and Lasso penalties.

It is useful when you think both zeroing out features and shrinking are helpful.

# ElasticNet regression
from sklearn.linear_model import ElasticNet
elastic = ElasticNet(alpha=1.0, l1_ratio=0.5)
elastic.fit(X_train, y_train)
elastic_train = elastic.score(X_train, y_train)
elastic_test = elastic.score(X_test, y_test)
print("ElasticNet Train R^2:", round(elastic_train, 3), "ElasticNet Test R^2:", round(elastic_test, 3))
ElasticNet Train R^2: 0.427 ElasticNet Test R^2: 0.417

Comparing Model Coefficients#

Let us look at the coefficients for each model.

The bigger the coefficient, the more that feature matters in the prediction.

# Compare the model coefficients side by side
coef_df = pd.DataFrame({
    "Feature": X.columns,
    "Linear": linreg.coef_,
    "Ridge": ridge.coef_,
    "Lasso": lasso.coef_,
    "ElasticNet": elastic.coef_
})
coef_df
Feature Linear Ridge Lasso ElasticNet
0 MedInc 0.448675 0.448511 0.148196 0.255275
1 HouseAge 0.009724 0.009726 0.005728 0.011230
2 AveRooms -0.123323 -0.123014 0.000000 0.000000
3 AveBedrms 0.783145 0.781417 -0.000000 -0.000000
4 Population -0.000002 -0.000002 -0.000008 0.000008
5 AveOccup -0.003526 -0.003526 -0.000000 -0.000000
6 Latitude -0.419792 -0.419787 -0.000000 -0.000000
7 Longitude -0.433708 -0.433681 -0.000000 -0.000000
# Practice: Which feature do you think is the most important for predicting house value?
answer = input("Type your guess from the list above: ")
print("Thanks for guessing! Let's look for it in the next cells.")
Thanks for guessing! Let's look for it in the next cells.
 

Best Practices: Choosing Alpha#

Picking the right alpha is important.

If alpha is too high, the model may underfit. If it is too low, the model may overfit.

Try different values and use validation data or cross-validation to choose the best one.

# Simple test: Try a smaller alpha with Ridge
ridge2 = Ridge(alpha=0.1)
ridge2.fit(X_train, y_train)
print("Ridge alpha=0.1 Test R^2:", round(ridge2.score(X_test, y_test), 3))
Ridge alpha=0.1 Test R^2: 0.576
# Tip: use cross-validation for the best alpha
from sklearn.model_selection import cross_val_score
scores = cross_val_score(Ridge(alpha=1.0), X_train, y_train, cv=5)
print("Average CV R^2:", round(scores.mean(), 3))
Average CV R^2: 0.611

Mini Project: Predicting with Regularized Models#

Let us build a mini-project. You will train a Ridge and a Lasso model, then predict values for the test set.

You will compare their mean squared error (MSE) to see which one is better for this data.

# Mini Project step 1: Fit both Ridge and Lasso, then predict on test set
ridge_full = Ridge(alpha=1.0)
lasso_full = Lasso(alpha=1.0)
ridge_full.fit(X_train, y_train)
lasso_full.fit(X_train, y_train)
ridge_pred = ridge_full.predict(X_test)
lasso_pred = lasso_full.predict(X_test)
# Mini Project step 2: Calculate and compare mean squared error (MSE)
from sklearn.metrics import mean_squared_error
ridge_mse = mean_squared_error(y_test, ridge_pred)
lasso_mse = mean_squared_error(y_test, lasso_pred)
print("Ridge MSE:", round(ridge_mse, 3))
print("Lasso MSE:", round(lasso_mse, 3))
Ridge MSE: 0.556
Lasso MSE: 0.938

Challenge: Try ElasticNet with a Custom l1_ratio#

Set the l1_ratio to 0.8 and see how the ElasticNet model error compares to Ridge and Lasso.

Which one is best for mean squared error?

Recap#

  • Regularization helps avoid overfitting by making models simpler.
  • Ridge keeps all features but shrinks them.
  • Lasso can remove useless features by making their weights zero.
  • ElasticNet combines both.

Remember, alpha is your tuning knob. Test different values and compare results!

Thanks for joining the lesson! For more fun Python tutorials, subscribe to our channel and check out the next videos.

Happy coding!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.