Mathew K Analytics

Lesson 4 · ML algorithms deep dive

Decision Trees: How They Really Work | ML Algorithms #4

Video four of the 12-part series: an algorithm that explains its own real reasoning. Real mushroom characteristics, labeled genuinely edible or genuinely…

What you'll learn

Datasets used in this lesson

Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.

📓 Full notebook

Download .ipynb

ML Algorithms Deep-Dive, Video 4: Decision Trees#

  • Video four of the 12-part series: an algorithm that explains its own real reasoning.
  • Real mushroom characteristics, labeled genuinely edible or genuinely poisonous.
  • Let's get into it.

Before You Start#

  • Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
  • You'll need pandas, NumPy, Matplotlib, and scikit-learn.
  • Place mushroom_dataset.csv in the same folder as this notebook.
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
from sklearn.tree import DecisionTreeClassifier, plot_tree, export_text
from sklearn.model_selection import train_test_split, cross_val_score, StratifiedKFold

mushrooms = pd.read_csv('mushroom_dataset.csv')
feature_cols = [c for c in mushrooms.columns if c not in ('actual', 'predicted')]
print(f'Real mushrooms in this dataset: {len(mushrooms)}')
print(f'Real features per mushroom: {len(feature_cols)}')
print(mushrooms['actual'].value_counts())
Real mushrooms in this dataset: 1625
Real features per mushroom: 22
actual
e    842
p    783
Name: count, dtype: int64

Part 1: The Core Idea#

X = mushrooms[feature_cols].values
y = mushrooms['actual'].values
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, random_state=42, stratify=y)
print(f'Real training mushrooms: {len(X_train)}')
print(f'Real test mushrooms: {len(X_test)}')
Real training mushrooms: 1218
Real test mushrooms: 407

Part 2: How a Split Gets Chosen#

def gini_impurity(labels):
    _, counts = np.unique(labels, return_counts=True)
    proportions = counts / counts.sum()
    return 1 - np.sum(proportions ** 2)

print(f'Real Gini impurity of the full real training set: {gini_impurity(y_train):.4f}')
edible_only = y_train[y_train == 'e']
print(f'Real Gini impurity of an all-edible real group: {gini_impurity(edible_only):.4f}')
Real Gini impurity of the full real training set: 0.4993
Real Gini impurity of an all-edible real group: 0.0000

Part 3: Fitting an Unconstrained Real Tree#

full_tree = DecisionTreeClassifier(random_state=42).fit(X_train, y_train)
print(f'Real training accuracy: {full_tree.score(X_train, y_train):.4f}')
print(f'Real test accuracy: {full_tree.score(X_test, y_test):.4f}')
print(f'Real tree depth: {full_tree.get_depth()}')
print(f'Real number of leaves: {full_tree.get_n_leaves()}')
Real training accuracy: 1.0000
Real test accuracy: 1.0000
Real tree depth: 8
Real number of leaves: 20

A real, honest surprise: both scores land at a perfect 1.0, and the unconstrained tree stopped naturally at just 8 real levels and 20 real leaves, not the sprawling, memorized structure a name like 'unconstrained' might suggest. This real dataset genuinely is that cleanly separable by a handful of physical mushroom traits; that's a property of the data, not something to expect from every real dataset.

Part 4: Controlling Complexity with Max Depth#

for depth in [1, 2, 3, 4, 5, None]:
    limited_tree = DecisionTreeClassifier(max_depth=depth, random_state=42).fit(X_train, y_train)
    train_acc = limited_tree.score(X_train, y_train)
    test_acc = limited_tree.score(X_test, y_test)
    print(f'max_depth={depth}: real train acc={train_acc:.4f}, real test acc={test_acc:.4f}')
max_depth=1: real train acc=0.8054, real test acc=0.7641
max_depth=2: real train acc=0.9048, real test acc=0.8993
max_depth=3: real train acc=0.9622, real test acc=0.9607
max_depth=4: real train acc=0.9836, real test acc=0.9779
max_depth=5: real train acc=0.9885, real test acc=0.9853
max_depth=None: real train acc=1.0000, real test acc=1.0000

Part 5: When Overfitting Actually Shows Up#

X_small, _, y_small, _ = train_test_split(X_train, y_train, train_size=60, random_state=42, stratify=y_train)
for depth in [1, 2, 3, None]:
    small_tree = DecisionTreeClassifier(max_depth=depth, random_state=42).fit(X_small, y_small)
    train_acc = small_tree.score(X_small, y_small)
    held_out_acc = small_tree.score(X_test, y_test)
    print(f'60-sample max_depth={depth}: real train acc={train_acc:.4f}, real held-out acc={held_out_acc:.4f}')
60-sample max_depth=1: real train acc=0.8167, real held-out acc=0.7273
60-sample max_depth=2: real train acc=0.9500, real held-out acc=0.8968
60-sample max_depth=3: real train acc=1.0000, real held-out acc=0.9435
60-sample max_depth=None: real train acc=1.0000, real held-out acc=0.9435

Even here, with a real training set shrunk to 60 mushrooms, deeper trees didn't overfit further, because the real underlying pattern is genuinely simple: a few real physical traits do almost all the work. A real dataset with messier, more overlapping classes would show a real accuracy gap opening up between train and held-out scores as depth grows; that gap is the real signal to watch for and prune against, whether or not it appears here.

Part 6: Reading the Tree#

shallow_tree = DecisionTreeClassifier(max_depth=3, random_state=42).fit(X_train, y_train)
print(export_text(shallow_tree, feature_names=feature_cols))
|--- gill-color <= 3.50
|   |--- population <= 3.50
|   |   |--- spore-print-color <= 1.50
|   |   |   |--- class: p
|   |   |--- spore-print-color >  1.50
|   |   |   |--- class: e
|   |--- population >  3.50
|   |   |--- stalk-root <= 1.00
|   |   |   |--- class: p
|   |   |--- stalk-root >  1.00
|   |   |   |--- class: e
|--- gill-color >  3.50
|   |--- spore-print-color <= 1.50
|   |   |--- odor <= 3.50
|   |   |   |--- class: p
|   |   |--- odor >  3.50
|   |   |   |--- class: e
|   |--- spore-print-color >  1.50
|   |   |--- gill-size <= 0.50
|   |   |   |--- class: e
|   |   |--- gill-size >  0.50
|   |   |   |--- class: p

plt.figure(figsize=(16, 8))
plot_tree(shallow_tree, feature_names=feature_cols, class_names=shallow_tree.classes_, filled=True, fontsize=8)
plt.title('Real Depth-3 Decision Tree')
plt.show()
No description has been provided for this image

Part 7: Feature Importance#

depth5_tree = DecisionTreeClassifier(max_depth=5, random_state=42).fit(X_train, y_train)
importances = pd.Series(depth5_tree.feature_importances_, index=feature_cols).sort_values(ascending=False)
print(importances.head(6))
gill-color           0.395064
population           0.176801
spore-print-color    0.162080
gill-size            0.143129
odor                 0.040705
stalk-shape          0.034263
dtype: float64

Part 8: Cross-Validating the Depth Choice#

skf = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
for depth in [1, 2, 3, 4, 5, 8, None]:
    scores = cross_val_score(DecisionTreeClassifier(max_depth=depth, random_state=42), X_train, y_train, cv=skf, scoring='accuracy')
    print(f'max_depth={depth}: real mean CV accuracy={scores.mean():.4f}')
max_depth=1: real mean CV accuracy=0.8054
max_depth=2: real mean CV accuracy=0.9047
max_depth=3: real mean CV accuracy=0.9622
max_depth=4: real mean CV accuracy=0.9803
max_depth=5: real mean CV accuracy=0.9828
max_depth=8: real mean CV accuracy=0.9943
max_depth=None: real mean CV accuracy=0.9943

Unlike video one's regularization sweep, where more complexity eventually hurt, this real search never found a real downside to growing further. That's not a contradiction, it's a reminder that the honest answer, 'how much complexity is too much,' genuinely depends on the real dataset in front of you, and has to be checked, not assumed either way.

Part 9: Strengths, Weaknesses, and When to Use It#

  • Strength: fully real, human-readable decision logic, no black box.
  • Strength: handles real categorical features natively, no scaling or encoding tricks required.
  • Weakness: a single tree can be unstable, small real changes in training data can produce a different real tree structure.
  • Weakness: prone to real overfitting on messier, noisier, or smaller real datasets than this one.
  • Use it when explaining the real 'why' behind a prediction matters as much as the prediction itself.

Part 10: Saving Your Work#

plt.figure(figsize=(18, 9))
plot_tree(depth5_tree, feature_names=feature_cols, class_names=depth5_tree.classes_, filled=True, fontsize=7)
plt.title('Real Depth-5 Decision Tree (Final Model)')
plt.savefig('decision_tree_final.png', dpi=150)
plt.show()
No description has been provided for this image
import joblib
joblib.dump(depth5_tree, 'decision_tree_model.joblib')
reloaded_tree = joblib.load('decision_tree_model.joblib')
print(f'Real original model test accuracy: {depth5_tree.score(X_test, y_test):.4f}')
print(f'Real reloaded model test accuracy: {reloaded_tree.score(X_test, y_test):.4f}')
Real original model test accuracy: 0.9853
Real reloaded model test accuracy: 0.9853

Wrap-Up: What You Learned#

  • Gini impurity, computed from scratch, and how it decides every real split.
  • An unconstrained real tree that turned out not to overfit, a genuine property of this particular real data.
  • A real overfitting gap that finally appeared once the training set was deliberately shrunk.
  • Reading a real tree's actual decision logic directly, in text and as a diagram.
  • Ranking real features by importance, and cross-validating the depth choice honestly.
  • Video five moves to random forests, which fix a single tree's real instability by growing many of them. Subscribe so it lands automatically see you there.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.