Mathew K Analytics

Lesson 11 · Scikit-learn deep dive

Scikit-learn Tutorial #11: Regression Metrics

Video eleven of the eighteen-part series: judging how far off a regressor's predictions really are. MSE, RMSE, MAE, R-squared, and explained variance. Let's…

⬇ Download notebookOpen in Colab ↗

📓 Full notebook

Download .ipynb

Scikit-learn Deep-Dive, Video 11: Regression Metrics#

  • Video eleven of the eighteen-part series: judging how far off a regressor's predictions really are.
  • MSE, RMSE, MAE, R-squared, and explained variance.
  • Let's get into it.

Part 1: mean_squared_error Basics#

import numpy as np
from sklearn.metrics import mean_squared_error
y_true = np.array([100, 150, 200, 250, 300])
y_pred = np.array([110, 140, 210, 230, 320])
mse = mean_squared_error(y_true, y_pred)
print(round(mse, 2))
220.0

Part 2: RMSE - Same Units as the Target#

rmse = np.sqrt(mse)
print(round(rmse, 2))
from sklearn.metrics import root_mean_squared_error
rmse_direct = root_mean_squared_error(y_true, y_pred)
print(round(rmse_direct, 2))
14.83
14.83

Part 3: mean_absolute_error - Robust to Outliers#

from sklearn.metrics import mean_absolute_error
mae = mean_absolute_error(y_true, y_pred)
print(round(mae, 2))
14.0

Part 4: Comparing MSE vs MAE - Sensitivity to an Outlier#

y_pred_outlier = np.array([110, 140, 210, 230, 500])
rmse_outlier = np.sqrt(mean_squared_error(y_true, y_pred_outlier))
mae_outlier = mean_absolute_error(y_true, y_pred_outlier)
print('RMSE change:', round(rmse_outlier - rmse, 2))
print('MAE change:', round(mae_outlier - mae, 2))
RMSE change: 75.39
MAE change: 36.0

Part 5: r2_score - Proportion of Variance Explained#

from sklearn.datasets import load_diabetes
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LinearRegression
from sklearn.metrics import r2_score
Xd, yd = load_diabetes(return_X_y=True)
Xd_train, Xd_test, yd_train, yd_test = train_test_split(Xd, yd, test_size=0.25, random_state=42)
reg = LinearRegression().fit(Xd_train, yd_train)
preds = reg.predict(Xd_test)
print(round(r2_score(yd_test, preds), 3))
0.485

Part 6: r2_score Can Be Negative - Worse Than Guessing the Mean#

bad_preds = np.full_like(yd_test, fill_value=yd_train.mean() * 3, dtype=float)
print(round(r2_score(yd_test, bad_preds), 3))
mean_baseline = np.full_like(yd_test, fill_value=yd_train.mean(), dtype=float)
print(round(r2_score(yd_test, mean_baseline), 3))
-18.229
-0.014

Part 7: explained_variance_score vs r2_score - When There's Bias#

from sklearn.metrics import explained_variance_score
biased_preds = preds + 20
print(round(r2_score(yd_test, biased_preds), 3))
print(round(explained_variance_score(yd_test, biased_preds), 3))
0.426
0.486

Part 8: mean_absolute_percentage_error#

from sklearn.metrics import mean_absolute_percentage_error
mape = mean_absolute_percentage_error(y_true, y_pred)
print(round(mape * 100, 2))
7.27

Part 9: Comparing Multiple Models with Several Metrics at Once#

from sklearn.tree import DecisionTreeRegressor
from sklearn.ensemble import RandomForestRegressor
models = {
    'linear': LinearRegression(),
    'tree': DecisionTreeRegressor(random_state=42, max_depth=4),
    'forest': RandomForestRegressor(random_state=42, n_estimators=100)
}
for name, m in models.items():
    m.fit(Xd_train, yd_train)
    p = m.predict(Xd_test)
    print(name, round(np.sqrt(mean_squared_error(yd_test, p)), 2), round(mean_absolute_error(yd_test, p), 2), round(r2_score(yd_test, p), 3))
linear 53.37 41.55 0.485
tree 59.65 46.77 0.357
forest 54.86 43.74 0.456

Part 10: A Real Pattern - a Reusable evaluate_regressor Function#

def evaluate_regressor(y_true, y_pred):
    return {
        'rmse': round(np.sqrt(mean_squared_error(y_true, y_pred)), 2),
        'mae': round(mean_absolute_error(y_true, y_pred), 2),
        'r2': round(r2_score(y_true, y_pred), 3)
    }
forest_preds = models['forest'].predict(Xd_test)
print(evaluate_regressor(yd_test, forest_preds))
{'rmse': np.float64(54.86), 'mae': 43.74, 'r2': 0.456}

Wrap-Up: What You Learned#

  • mean_squared_error averages squared differences, disproportionately punishing larger errors.
  • RMSE is the square root of MSE, back in the original units of the target and directly interpretable.
  • mean_absolute_error averages plain absolute differences and is more robust to outliers than MSE or RMSE.
  • A single large outlier moves RMSE far more than it moves MAE.
  • r2_score reports the proportion of variance explained; one is perfect, zero matches always guessing the mean.
  • r2_score can go negative, meaning a model performs worse than the trivial mean-guessing baseline.
  • explained_variance_score ignores a constant systematic bias in predictions, while r2_score still penalizes it.
  • mean_absolute_percentage_error reports error as a percentage, useful for comparing across differently-scaled targets.
  • Comparing several metrics across candidate models at once reveals different error patterns a single number would hide.
  • That wraps up regression metrics. Next up: Clustering Metrics - silhouette score and the elbow method.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.