Lesson 5 · ML algorithms deep dive
Random Forests Explained with Python | ML Algorithms #5
Video five of the 12-part series: fixing a single tree's instability by growing many of them. Real chemical measurements from 178 real wines, across three…
- CourseML algorithms deep dive
- Lesson5 of 12
- Video17 min
- FormatJupyter notebook · 12 code cells
- Data1 dataset
What you'll learn
Datasets used in this lesson
Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.
- wine.csv12.0 KB
📓 Full notebook
Download .ipynbML Algorithms Deep-Dive, Video 5: Random Forests#
- Video five of the 12-part series: fixing a single tree's instability by growing many of them.
- Real chemical measurements from 178 real wines, across three real cultivars.
- Let's get into it.
Before You Start#
- Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
- You'll need pandas, NumPy, Matplotlib, SciPy, and scikit-learn.
- Place wine.csv in the same folder as this notebook.
import pandas as pd
import numpy as np
from scipy import stats
from sklearn.tree import DecisionTreeClassifier
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import train_test_split, cross_val_score, StratifiedKFold
wine = pd.read_csv('wine.csv')
feature_cols = [c for c in wine.columns if c != 'target']
print(f'Real wines in this dataset: {len(wine)}')
print(f'Real chemical features per wine: {len(feature_cols)}')
print(wine['target'].value_counts())
Part 1: The Core Idea#
X = wine[feature_cols].values
y = wine['target'].values
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, random_state=42, stratify=y)
print(f'Real training wines: {len(X_train)}')
print(f'Real test wines: {len(X_test)}')
Part 2: Single-Tree Instability, Demonstrated#
rng = np.random.RandomState(42)
n_train = len(X_train)
all_bootstrap_preds = []
for i in range(20):
sample_idx = rng.choice(n_train, size=n_train, replace=True)
tree = DecisionTreeClassifier(random_state=i).fit(X_train[sample_idx], y_train[sample_idx])
all_bootstrap_preds.append(tree.predict(X_test))
all_bootstrap_preds = np.array(all_bootstrap_preds)
print(f'Real shape of the prediction grid: {all_bootstrap_preds.shape}')
disagreement_count = 0
for col in range(all_bootstrap_preds.shape[1]):
if len(set(all_bootstrap_preds[:, col])) > 1:
disagreement_count += 1
print(f'Real test wines where the 20 bootstrap trees disagreed: {disagreement_count} of {all_bootstrap_preds.shape[1]}')
Part 3: Bagging Fixes It by Voting#
majority_vote = stats.mode(all_bootstrap_preds, axis=0, keepdims=False).mode
bagged_accuracy = (majority_vote == y_test).mean()
single_tree_accuracy = DecisionTreeClassifier(random_state=42).fit(X_train, y_train).score(X_test, y_test)
print(f'Real single-tree test accuracy: {single_tree_accuracy:.4f}')
print(f'Real 20-tree majority-vote accuracy: {bagged_accuracy:.4f}')
Part 4: Fitting a Real Random Forest#
for n_estimators in [1, 10, 50, 100, 200]:
forest = RandomForestClassifier(n_estimators=n_estimators, random_state=42).fit(X_train, y_train)
print(f'n_estimators={n_estimators}: real test accuracy={forest.score(X_test, y_test):.4f}')
Part 5: Out-of-Bag Score: Free Validation#
forest_oob = RandomForestClassifier(n_estimators=200, random_state=42, oob_score=True).fit(X_train, y_train)
print(f'Real out-of-bag score: {forest_oob.oob_score_:.4f}')
print(f'Real held-out test accuracy: {forest_oob.score(X_test, y_test):.4f}')
Part 6: Feature Randomness: The Second Real Ingredient#
skf = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
for max_features in ['sqrt', 'log2', None]:
scores = cross_val_score(RandomForestClassifier(n_estimators=100, max_features=max_features, random_state=42), X_train, y_train, cv=skf, scoring='accuracy')
print(f'max_features={max_features}: real mean CV accuracy={scores.mean():.4f}')
Part 7: Feature Importance, Spread Out#
forest_final = RandomForestClassifier(n_estimators=200, random_state=42).fit(X_train, y_train)
forest_importances = pd.Series(forest_final.feature_importances_, index=feature_cols).sort_values(ascending=False)
print('Real random forest feature importances:')
print(forest_importances.head(6))
single_tree_final = DecisionTreeClassifier(random_state=42).fit(X_train, y_train)
tree_importances = pd.Series(single_tree_final.feature_importances_, index=feature_cols).sort_values(ascending=False)
print('Real single-tree feature importances:')
print(tree_importances.head(6))
That contrast is the real, visible fingerprint of feature randomness from part six. Because every real tree in the forest is occasionally forced to split on something other than the single best real feature, the forest as a whole ends up genuinely using more of the real available information, not just leaning on the same one or two features a single tree would default to every time.
Part 8: How Many Trees Are Actually Needed?#
for n_estimators in [10, 50, 100, 200]:
scores = cross_val_score(RandomForestClassifier(n_estimators=n_estimators, random_state=42), X_train, y_train, cv=skf, scoring='accuracy')
print(f'n_estimators={n_estimators}: real CV mean={scores.mean():.4f}, real CV std={scores.std():.4f}')
Part 9: Strengths, Weaknesses, and When to Use It#
- Strength: consistently more real accurate and more real stable than a single tree, as parts two and three showed directly.
- Strength: a genuinely free, built-in validation estimate via the real out-of-bag score.
- Weakness: loses the single tree's real, plain-English explainability, a forest of 200 trees can't be read like the diagram in video four.
- Weakness: more real compute and real memory than a single tree, for real but often modest accuracy gains past a certain size.
- Use it as a genuinely strong real default for tabular classification whenever a single tree's instability, from part two, is a real concern.
Part 10: Saving Your Work#
import matplotlib.pyplot as plt
plt.figure(figsize=(9, 6))
forest_importances.head(10).sort_values().plot(kind='barh', color='seagreen')
plt.title('Real Random Forest Feature Importances (Top 10)')
plt.xlabel('Importance')
plt.savefig('random_forest_feature_importance.png', dpi=150)
plt.show()
import joblib
joblib.dump(forest_final, 'random_forest_model.joblib')
reloaded_forest = joblib.load('random_forest_model.joblib')
print(f'Real original model test accuracy: {forest_final.score(X_test, y_test):.4f}')
print(f'Real reloaded model test accuracy: {reloaded_forest.score(X_test, y_test):.4f}')
Wrap-Up: What You Learned#
- A real, direct demonstration of single-tree instability, over a third of real test wines saw genuine disagreement across bootstrap-resampled trees.
- Bagging: a real majority vote across many trees, fixing that instability.
- The out-of-bag score, a genuinely free validation estimate baked into every random forest.
- Feature randomness, the second real ingredient beyond bagging, and real proof it adds value beyond resampling alone.
- Why forest feature importance spreads out more evenly than a single tree's.
- Cross-validating forest size to find where real accuracy gains actually stop.
- Video six moves to gradient boosting, a completely different way to combine many real trees, on a real regression problem. Subscribe so it lands automatically see you there.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



