Lesson 6 · ML algorithms deep dive
Gradient Boosting Explained Step by Step | ML Algorithms #6
Video six of the 12-part series: a genuinely different way to combine many real trees. Real patient measurements predicting real diabetes disease…
- CourseML algorithms deep dive
- Lesson6 of 12
- Video17 min
- FormatJupyter notebook · 12 code cells
- Data1 dataset
What you'll learn
Datasets used in this lesson
Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.
- diabetes.csv92.9 KB
📓 Full notebook
Download .ipynbML Algorithms Deep-Dive, Video 6: Gradient Boosting#
- Video six of the 12-part series: a genuinely different way to combine many real trees.
- Real patient measurements predicting real diabetes disease progression.
- Let's get into it.
Before You Start#
- Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
- You'll need pandas, NumPy, Matplotlib, and scikit-learn.
- Place diabetes.csv in the same folder as this notebook.
import pandas as pd
import numpy as np
from sklearn.linear_model import LinearRegression
from sklearn.tree import DecisionTreeRegressor
from sklearn.ensemble import GradientBoostingRegressor
from sklearn.model_selection import train_test_split, cross_val_score, KFold
from sklearn.metrics import mean_absolute_error, r2_score
diabetes = pd.read_csv('diabetes.csv')
print(f'Real patients in this dataset: {len(diabetes)}')
print(diabetes[["bmi", "bp", "target"]].describe())
Part 1: A Genuinely Different Way to Combine Trees#
X = diabetes.drop(columns=['target']).values
y = diabetes['target'].values
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, random_state=42)
print(f'Real training patients: {len(X_train)}')
print(f'Real test patients: {len(X_test)}')
Part 2: An Honest Baseline#
linear_model = LinearRegression().fit(X_train, y_train)
linear_pred = linear_model.predict(X_test)
print(f'Real linear regression test R-squared: {r2_score(y_test, linear_pred):.4f}')
print(f'Real linear regression test MAE: {mean_absolute_error(y_test, linear_pred):.4f}')
Part 3: Boosting From Scratch, Fitting the Residuals#
tree_1 = DecisionTreeRegressor(max_depth=2, random_state=42).fit(X_train, y_train)
pred_1 = tree_1.predict(X_train)
residual_1 = y_train - pred_1
print(f'Real training MAE after tree 1: {np.abs(residual_1).mean():.4f}')
learning_rate = 0.1
tree_2 = DecisionTreeRegressor(max_depth=2, random_state=42).fit(X_train, residual_1)
combined_pred = pred_1 + learning_rate * tree_2.predict(X_train)
residual_2 = y_train - combined_pred
print(f'Real training MAE after tree 2: {np.abs(residual_2).mean():.4f}')
tree_3 = DecisionTreeRegressor(max_depth=2, random_state=42).fit(X_train, residual_2)
combined_pred = combined_pred + learning_rate * tree_3.predict(X_train)
residual_3 = y_train - combined_pred
print(f'Real training MAE after tree 3: {np.abs(residual_3).mean():.4f}')
Part 4: Fitting a Real Gradient Boosting Model#
gbr_default = GradientBoostingRegressor(random_state=42).fit(X_train, y_train)
gbr_pred = gbr_default.predict(X_test)
print(f'Real default gradient boosting test R-squared: {r2_score(y_test, gbr_pred):.4f}')
print(f'Real default gradient boosting test MAE: {mean_absolute_error(y_test, gbr_pred):.4f}')
Part 5: The Learning Rate and Tree Count Tradeoff#
for rate in [0.01, 0.1, 0.3]:
for n_trees in [50, 100, 300]:
model = GradientBoostingRegressor(learning_rate=rate, n_estimators=n_trees, random_state=42).fit(X_train, y_train)
preds = model.predict(X_test)
print(f'rate={rate}, n_trees={n_trees}: real test R-squared={r2_score(y_test, preds):.4f}, real MAE={mean_absolute_error(y_test, preds):.4f}')
That's real overfitting, directly visible: a large real learning rate combined with many real trees lets the model chase the real training data's noise too aggressively. A small real learning rate needs more real trees to reach the same place, but gets there more carefully, and this real search's single best combination, rate 0.01 with 300 trees, still only reaches an R-squared of 0.47, just short of part two's plain linear regression.
Part 6: Confirming It With Cross-Validation#
kfold = KFold(n_splits=5, shuffle=True, random_state=42)
linear_cv = cross_val_score(LinearRegression(), X_train, y_train, cv=kfold, scoring='r2')
gbr_cv = cross_val_score(GradientBoostingRegressor(learning_rate=0.01, n_estimators=300, random_state=42), X_train, y_train, cv=kfold, scoring='r2')
print(f'Real linear regression mean CV R-squared: {linear_cv.mean():.4f}')
print(f'Real best gradient boosting mean CV R-squared: {gbr_cv.mean():.4f}')
Diabetes progression, from these ten real baseline measurements, is apparently close enough to a real linear relationship that a far more flexible, sequentially-corrected model can't find genuine extra signal to exploit, it mostly finds real noise to overfit instead. That's an honest, valuable real finding: real algorithm sophistication only helps when there's real non-linear structure worth capturing.
Part 7: Early Stopping#
gbr_early = GradientBoostingRegressor(n_estimators=1000, learning_rate=0.05, validation_fraction=0.2, n_iter_no_change=10, random_state=42).fit(X_train, y_train)
print(f'Real trees actually used before stopping: {gbr_early.n_estimators_}')
early_pred = gbr_early.predict(X_test)
print(f'Real early-stopped test R-squared: {r2_score(y_test, early_pred):.4f}')
Part 8: Feature Importance#
gbr_best = GradientBoostingRegressor(learning_rate=0.01, n_estimators=300, random_state=42).fit(X_train, y_train)
feature_names = diabetes.drop(columns=['target']).columns
importances = pd.Series(gbr_best.feature_importances_, index=feature_names).sort_values(ascending=False)
print(importances)
Part 9: Strengths, Weaknesses, and When to Use It#
- Strength: often the strongest real performer on structured, tabular data with genuine non-linear patterns.
- Strength: early stopping, from part seven, automates a real chunk of hyperparameter tuning.
- Weakness: sensitive to real learning rate and tree count together, as part five showed directly.
- Weakness: not automatically better than a real simpler model, as this entire real dataset just demonstrated.
- Use it when a real baseline like linear regression or a single tree has already been tried and genuinely underperforms.
Part 10: Saving Your Work#
import matplotlib.pyplot as plt
plt.figure(figsize=(9, 6))
importances.sort_values().plot(kind='barh', color='darkorange')
plt.title('Real Gradient Boosting Feature Importances')
plt.xlabel('Importance')
plt.savefig('gradient_boosting_feature_importance.png', dpi=150)
plt.show()
import joblib
joblib.dump(gbr_best, 'gradient_boosting_model.joblib')
reloaded_gbr = joblib.load('gradient_boosting_model.joblib')
print(f'Real original model test R-squared: {gbr_best.score(X_test, y_test):.4f}')
print(f'Real reloaded model test R-squared: {reloaded_gbr.score(X_test, y_test):.4f}')
Wrap-Up: What You Learned#
- Gradient boosting's real core idea: each new tree corrects the real errors of everything fit before it, built from scratch across three real rounds.
- Fitting scikit-learn's real gradient boosting regressor, and an honest result where it lost to plain linear regression.
- The real learning rate and tree count tradeoff, with overfitting directly visible in a 3 by 3 real sweep.
- Confirming that result with real cross-validation, not just a single lucky or unlucky split.
- Early stopping, letting the real model choose its own tree count automatically.
- A real, honest reminder that a more sophisticated algorithm only wins when there's genuine structure for it to find.
- Video seven moves to support vector machines, revisiting real breast cancer data for a direct comparison against video two's logistic regression. Subscribe so it lands automatically see you there.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



