Mathew K Analytics

Lesson 6 · ML algorithms deep dive

Gradient Boosting Explained Step by Step | ML Algorithms #6

Video six of the 12-part series: a genuinely different way to combine many real trees. Real patient measurements predicting real diabetes disease…

What you'll learn

Datasets used in this lesson

Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.

📓 Full notebook

Download .ipynb

ML Algorithms Deep-Dive, Video 6: Gradient Boosting#

  • Video six of the 12-part series: a genuinely different way to combine many real trees.
  • Real patient measurements predicting real diabetes disease progression.
  • Let's get into it.

Before You Start#

  • Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
  • You'll need pandas, NumPy, Matplotlib, and scikit-learn.
  • Place diabetes.csv in the same folder as this notebook.
import pandas as pd
import numpy as np
from sklearn.linear_model import LinearRegression
from sklearn.tree import DecisionTreeRegressor
from sklearn.ensemble import GradientBoostingRegressor
from sklearn.model_selection import train_test_split, cross_val_score, KFold
from sklearn.metrics import mean_absolute_error, r2_score

diabetes = pd.read_csv('diabetes.csv')
print(f'Real patients in this dataset: {len(diabetes)}')
print(diabetes[["bmi", "bp", "target"]].describe())
Real patients in this dataset: 442
                bmi            bp      target
count  4.420000e+02  4.420000e+02  442.000000
mean  -2.087320e-16 -4.571507e-17  152.133484
std    4.761905e-02  4.761905e-02   77.093005
min   -9.027530e-02 -1.123988e-01   25.000000
25%   -3.422907e-02 -3.665608e-02   87.000000
50%   -7.283766e-03 -5.670422e-03  140.500000
75%    3.124802e-02  3.564379e-02  211.500000
max    1.705552e-01  1.320436e-01  346.000000

Part 1: A Genuinely Different Way to Combine Trees#

X = diabetes.drop(columns=['target']).values
y = diabetes['target'].values
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, random_state=42)
print(f'Real training patients: {len(X_train)}')
print(f'Real test patients: {len(X_test)}')
Real training patients: 331
Real test patients: 111

Part 2: An Honest Baseline#

linear_model = LinearRegression().fit(X_train, y_train)
linear_pred = linear_model.predict(X_test)
print(f'Real linear regression test R-squared: {r2_score(y_test, linear_pred):.4f}')
print(f'Real linear regression test MAE: {mean_absolute_error(y_test, linear_pred):.4f}')
Real linear regression test R-squared: 0.4849
Real linear regression test MAE: 41.5485

Part 3: Boosting From Scratch, Fitting the Residuals#

tree_1 = DecisionTreeRegressor(max_depth=2, random_state=42).fit(X_train, y_train)
pred_1 = tree_1.predict(X_train)
residual_1 = y_train - pred_1
print(f'Real training MAE after tree 1: {np.abs(residual_1).mean():.4f}')
Real training MAE after tree 1: 47.9960
learning_rate = 0.1
tree_2 = DecisionTreeRegressor(max_depth=2, random_state=42).fit(X_train, residual_1)
combined_pred = pred_1 + learning_rate * tree_2.predict(X_train)
residual_2 = y_train - combined_pred
print(f'Real training MAE after tree 2: {np.abs(residual_2).mean():.4f}')
tree_3 = DecisionTreeRegressor(max_depth=2, random_state=42).fit(X_train, residual_2)
combined_pred = combined_pred + learning_rate * tree_3.predict(X_train)
residual_3 = y_train - combined_pred
print(f'Real training MAE after tree 3: {np.abs(residual_3).mean():.4f}')
Real training MAE after tree 2: 47.4083
Real training MAE after tree 3: 46.9077

Part 4: Fitting a Real Gradient Boosting Model#

gbr_default = GradientBoostingRegressor(random_state=42).fit(X_train, y_train)
gbr_pred = gbr_default.predict(X_test)
print(f'Real default gradient boosting test R-squared: {r2_score(y_test, gbr_pred):.4f}')
print(f'Real default gradient boosting test MAE: {mean_absolute_error(y_test, gbr_pred):.4f}')
Real default gradient boosting test R-squared: 0.4241
Real default gradient boosting test MAE: 44.2497

Part 5: The Learning Rate and Tree Count Tradeoff#

for rate in [0.01, 0.1, 0.3]:
    for n_trees in [50, 100, 300]:
        model = GradientBoostingRegressor(learning_rate=rate, n_estimators=n_trees, random_state=42).fit(X_train, y_train)
        preds = model.predict(X_test)
        print(f'rate={rate}, n_trees={n_trees}: real test R-squared={r2_score(y_test, preds):.4f}, real MAE={mean_absolute_error(y_test, preds):.4f}')
rate=0.01, n_trees=50: real test R-squared=0.2817, real MAE=54.3286
rate=0.01, n_trees=100: real test R-squared=0.4068, real MAE=49.0545
rate=0.01, n_trees=300: real test R-squared=0.4702, real MAE=43.4553
rate=0.1, n_trees=50: real test R-squared=0.4418, real MAE=43.9628
rate=0.1, n_trees=100: real test R-squared=0.4241, real MAE=44.2497
rate=0.1, n_trees=300: real test R-squared=0.3765, real MAE=46.2551
rate=0.3, n_trees=50: real test R-squared=0.4269, real MAE=44.1177
rate=0.3, n_trees=100: real test R-squared=0.3825, real MAE=45.9672
rate=0.3, n_trees=300: real test R-squared=0.3211, real MAE=47.6979

That's real overfitting, directly visible: a large real learning rate combined with many real trees lets the model chase the real training data's noise too aggressively. A small real learning rate needs more real trees to reach the same place, but gets there more carefully, and this real search's single best combination, rate 0.01 with 300 trees, still only reaches an R-squared of 0.47, just short of part two's plain linear regression.

Part 6: Confirming It With Cross-Validation#

kfold = KFold(n_splits=5, shuffle=True, random_state=42)
linear_cv = cross_val_score(LinearRegression(), X_train, y_train, cv=kfold, scoring='r2')
gbr_cv = cross_val_score(GradientBoostingRegressor(learning_rate=0.01, n_estimators=300, random_state=42), X_train, y_train, cv=kfold, scoring='r2')
print(f'Real linear regression mean CV R-squared: {linear_cv.mean():.4f}')
print(f'Real best gradient boosting mean CV R-squared: {gbr_cv.mean():.4f}')
Real linear regression mean CV R-squared: 0.4606
Real best gradient boosting mean CV R-squared: 0.3426

Diabetes progression, from these ten real baseline measurements, is apparently close enough to a real linear relationship that a far more flexible, sequentially-corrected model can't find genuine extra signal to exploit, it mostly finds real noise to overfit instead. That's an honest, valuable real finding: real algorithm sophistication only helps when there's real non-linear structure worth capturing.

Part 7: Early Stopping#

gbr_early = GradientBoostingRegressor(n_estimators=1000, learning_rate=0.05, validation_fraction=0.2, n_iter_no_change=10, random_state=42).fit(X_train, y_train)
print(f'Real trees actually used before stopping: {gbr_early.n_estimators_}')
early_pred = gbr_early.predict(X_test)
print(f'Real early-stopped test R-squared: {r2_score(y_test, early_pred):.4f}')
Real trees actually used before stopping: 43
Real early-stopped test R-squared: 0.4570

Part 8: Feature Importance#

gbr_best = GradientBoostingRegressor(learning_rate=0.01, n_estimators=300, random_state=42).fit(X_train, y_train)
feature_names = diabetes.drop(columns=['target']).columns
importances = pd.Series(gbr_best.feature_importances_, index=feature_names).sort_values(ascending=False)
print(importances)
bmi    0.429919
s5     0.240100
bp     0.100489
s6     0.039551
age    0.038346
s3     0.037029
s1     0.035839
s2     0.035485
s4     0.033014
sex    0.010228
dtype: float64

Part 9: Strengths, Weaknesses, and When to Use It#

  • Strength: often the strongest real performer on structured, tabular data with genuine non-linear patterns.
  • Strength: early stopping, from part seven, automates a real chunk of hyperparameter tuning.
  • Weakness: sensitive to real learning rate and tree count together, as part five showed directly.
  • Weakness: not automatically better than a real simpler model, as this entire real dataset just demonstrated.
  • Use it when a real baseline like linear regression or a single tree has already been tried and genuinely underperforms.

Part 10: Saving Your Work#

import matplotlib.pyplot as plt
plt.figure(figsize=(9, 6))
importances.sort_values().plot(kind='barh', color='darkorange')
plt.title('Real Gradient Boosting Feature Importances')
plt.xlabel('Importance')
plt.savefig('gradient_boosting_feature_importance.png', dpi=150)
plt.show()
No description has been provided for this image
import joblib
joblib.dump(gbr_best, 'gradient_boosting_model.joblib')
reloaded_gbr = joblib.load('gradient_boosting_model.joblib')
print(f'Real original model test R-squared: {gbr_best.score(X_test, y_test):.4f}')
print(f'Real reloaded model test R-squared: {reloaded_gbr.score(X_test, y_test):.4f}')
Real original model test R-squared: 0.4702
Real reloaded model test R-squared: 0.4702

Wrap-Up: What You Learned#

  • Gradient boosting's real core idea: each new tree corrects the real errors of everything fit before it, built from scratch across three real rounds.
  • Fitting scikit-learn's real gradient boosting regressor, and an honest result where it lost to plain linear regression.
  • The real learning rate and tree count tradeoff, with overfitting directly visible in a 3 by 3 real sweep.
  • Confirming that result with real cross-validation, not just a single lucky or unlucky split.
  • Early stopping, letting the real model choose its own tree count automatically.
  • A real, honest reminder that a more sophisticated algorithm only wins when there's genuine structure for it to find.
  • Video seven moves to support vector machines, revisiting real breast cancer data for a direct comparison against video two's logistic regression. Subscribe so it lands automatically see you there.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.