Mathew K Analytics

Lesson 5 · ML algorithms deep dive

Random Forests Explained with Python | ML Algorithms #5

Video five of the 12-part series: fixing a single tree's instability by growing many of them. Real chemical measurements from 178 real wines, across three…

What you'll learn

Datasets used in this lesson

Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.

📓 Full notebook

Download .ipynb

ML Algorithms Deep-Dive, Video 5: Random Forests#

  • Video five of the 12-part series: fixing a single tree's instability by growing many of them.
  • Real chemical measurements from 178 real wines, across three real cultivars.
  • Let's get into it.

Before You Start#

  • Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
  • You'll need pandas, NumPy, Matplotlib, SciPy, and scikit-learn.
  • Place wine.csv in the same folder as this notebook.
import pandas as pd
import numpy as np
from scipy import stats
from sklearn.tree import DecisionTreeClassifier
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import train_test_split, cross_val_score, StratifiedKFold

wine = pd.read_csv('wine.csv')
feature_cols = [c for c in wine.columns if c != 'target']
print(f'Real wines in this dataset: {len(wine)}')
print(f'Real chemical features per wine: {len(feature_cols)}')
print(wine['target'].value_counts())
Real wines in this dataset: 178
Real chemical features per wine: 13
target
1    71
0    59
2    48
Name: count, dtype: int64

Part 1: The Core Idea#

X = wine[feature_cols].values
y = wine['target'].values
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, random_state=42, stratify=y)
print(f'Real training wines: {len(X_train)}')
print(f'Real test wines: {len(X_test)}')
Real training wines: 133
Real test wines: 45

Part 2: Single-Tree Instability, Demonstrated#

rng = np.random.RandomState(42)
n_train = len(X_train)
all_bootstrap_preds = []
for i in range(20):
    sample_idx = rng.choice(n_train, size=n_train, replace=True)
    tree = DecisionTreeClassifier(random_state=i).fit(X_train[sample_idx], y_train[sample_idx])
    all_bootstrap_preds.append(tree.predict(X_test))
all_bootstrap_preds = np.array(all_bootstrap_preds)
print(f'Real shape of the prediction grid: {all_bootstrap_preds.shape}')
Real shape of the prediction grid: (20, 45)
disagreement_count = 0
for col in range(all_bootstrap_preds.shape[1]):
    if len(set(all_bootstrap_preds[:, col])) > 1:
        disagreement_count += 1
print(f'Real test wines where the 20 bootstrap trees disagreed: {disagreement_count} of {all_bootstrap_preds.shape[1]}')
Real test wines where the 20 bootstrap trees disagreed: 16 of 45

Part 3: Bagging Fixes It by Voting#

majority_vote = stats.mode(all_bootstrap_preds, axis=0, keepdims=False).mode
bagged_accuracy = (majority_vote == y_test).mean()
single_tree_accuracy = DecisionTreeClassifier(random_state=42).fit(X_train, y_train).score(X_test, y_test)
print(f'Real single-tree test accuracy: {single_tree_accuracy:.4f}')
print(f'Real 20-tree majority-vote accuracy: {bagged_accuracy:.4f}')
Real single-tree test accuracy: 0.9556
Real 20-tree majority-vote accuracy: 1.0000

Part 4: Fitting a Real Random Forest#

for n_estimators in [1, 10, 50, 100, 200]:
    forest = RandomForestClassifier(n_estimators=n_estimators, random_state=42).fit(X_train, y_train)
    print(f'n_estimators={n_estimators}: real test accuracy={forest.score(X_test, y_test):.4f}')
n_estimators=1: real test accuracy=0.8667
n_estimators=10: real test accuracy=0.9778
n_estimators=50: real test accuracy=1.0000
n_estimators=100: real test accuracy=1.0000
n_estimators=200: real test accuracy=1.0000

Part 5: Out-of-Bag Score: Free Validation#

forest_oob = RandomForestClassifier(n_estimators=200, random_state=42, oob_score=True).fit(X_train, y_train)
print(f'Real out-of-bag score: {forest_oob.oob_score_:.4f}')
print(f'Real held-out test accuracy: {forest_oob.score(X_test, y_test):.4f}')
Real out-of-bag score: 0.9699
Real held-out test accuracy: 1.0000

Part 6: Feature Randomness: The Second Real Ingredient#

skf = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
for max_features in ['sqrt', 'log2', None]:
    scores = cross_val_score(RandomForestClassifier(n_estimators=100, max_features=max_features, random_state=42), X_train, y_train, cv=skf, scoring='accuracy')
    print(f'max_features={max_features}: real mean CV accuracy={scores.mean():.4f}')
max_features=sqrt: real mean CV accuracy=0.9852
max_features=log2: real mean CV accuracy=0.9852
max_features=None: real mean CV accuracy=0.9479

Part 7: Feature Importance, Spread Out#

forest_final = RandomForestClassifier(n_estimators=200, random_state=42).fit(X_train, y_train)
forest_importances = pd.Series(forest_final.feature_importances_, index=feature_cols).sort_values(ascending=False)
print('Real random forest feature importances:')
print(forest_importances.head(6))
single_tree_final = DecisionTreeClassifier(random_state=42).fit(X_train, y_train)
tree_importances = pd.Series(single_tree_final.feature_importances_, index=feature_cols).sort_values(ascending=False)
print('Real single-tree feature importances:')
print(tree_importances.head(6))
Real random forest feature importances:
color_intensity                 0.175995
flavanoids                      0.168809
proline                         0.151254
alcohol                         0.127567
hue                             0.095100
od280/od315_of_diluted_wines    0.088004
dtype: float64
Real single-tree feature importances:
flavanoids                      0.410802
color_intensity                 0.403317
proline                         0.100318
od280/od315_of_diluted_wines    0.022351
alcalinity_of_ash               0.022219
ash                             0.021419
dtype: float64

That contrast is the real, visible fingerprint of feature randomness from part six. Because every real tree in the forest is occasionally forced to split on something other than the single best real feature, the forest as a whole ends up genuinely using more of the real available information, not just leaning on the same one or two features a single tree would default to every time.

Part 8: How Many Trees Are Actually Needed?#

for n_estimators in [10, 50, 100, 200]:
    scores = cross_val_score(RandomForestClassifier(n_estimators=n_estimators, random_state=42), X_train, y_train, cv=skf, scoring='accuracy')
    print(f'n_estimators={n_estimators}: real CV mean={scores.mean():.4f}, real CV std={scores.std():.4f}')
n_estimators=10: real CV mean=0.9701, real CV std=0.0278
n_estimators=50: real CV mean=0.9852, real CV std=0.0296
n_estimators=100: real CV mean=0.9852, real CV std=0.0296
n_estimators=200: real CV mean=0.9852, real CV std=0.0296

Part 9: Strengths, Weaknesses, and When to Use It#

  • Strength: consistently more real accurate and more real stable than a single tree, as parts two and three showed directly.
  • Strength: a genuinely free, built-in validation estimate via the real out-of-bag score.
  • Weakness: loses the single tree's real, plain-English explainability, a forest of 200 trees can't be read like the diagram in video four.
  • Weakness: more real compute and real memory than a single tree, for real but often modest accuracy gains past a certain size.
  • Use it as a genuinely strong real default for tabular classification whenever a single tree's instability, from part two, is a real concern.

Part 10: Saving Your Work#

import matplotlib.pyplot as plt
plt.figure(figsize=(9, 6))
forest_importances.head(10).sort_values().plot(kind='barh', color='seagreen')
plt.title('Real Random Forest Feature Importances (Top 10)')
plt.xlabel('Importance')
plt.savefig('random_forest_feature_importance.png', dpi=150)
plt.show()
No description has been provided for this image
import joblib
joblib.dump(forest_final, 'random_forest_model.joblib')
reloaded_forest = joblib.load('random_forest_model.joblib')
print(f'Real original model test accuracy: {forest_final.score(X_test, y_test):.4f}')
print(f'Real reloaded model test accuracy: {reloaded_forest.score(X_test, y_test):.4f}')
Real original model test accuracy: 1.0000
Real reloaded model test accuracy: 1.0000

Wrap-Up: What You Learned#

  • A real, direct demonstration of single-tree instability, over a third of real test wines saw genuine disagreement across bootstrap-resampled trees.
  • Bagging: a real majority vote across many trees, fixing that instability.
  • The out-of-bag score, a genuinely free validation estimate baked into every random forest.
  • Feature randomness, the second real ingredient beyond bagging, and real proof it adds value beyond resampling alone.
  • Why forest feature importance spreads out more evenly than a single tree's.
  • Cross-validating forest size to find where real accuracy gains actually stop.
  • Video six moves to gradient boosting, a completely different way to combine many real trees, on a real regression problem. Subscribe so it lands automatically see you there.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.