Mathew K Analytics

Lesson 61 · Python for Data Science

Understanding Decision Trees and Random Forests in Python for Classification

Today, you will learn how computers make decisions by asking yes/no questions, using decision trees. You will also see how random forests combine many trees…

⬇ Download notebookOpen in Colab ↗

📓 Full notebook

Download .ipynb
 

Welcome to Decision Trees and Random Forests in Python!#

Today, you will learn how computers make decisions by asking yes/no questions, using decision trees.

You will also see how random forests combine many trees to make even smarter choices.

By the end, you will build and use models to help computers predict things in the real world!

What is a Decision Tree?#

A decision tree is a flowchart-like structure.

Each point, or node, asks a question.

The tree splits into branches based on answers.

They help computers pick or predict an answer, step by step.

# Let us build our first decision tree with scikit-learn, a popular Python library.
from sklearn import tree
X = [[0, 0], [1, 1]]  # feature data
y = [0, 1]             # labels: 0 or 1
classifier = tree.DecisionTreeClassifier()
classifier = classifier.fit(X, y)
print('Prediction for [2, 2]:', classifier.predict([[2, 2]]))
Prediction for [2, 2]: [1]

The Data: Features and Labels#

Features are facts or numbers we know about each example.

Labels are the outcomes we want to predict.

For example, features could be test scores; labels could be if a student passed or not.

# Here is a pretend dataset about students and passing status.
X = [[90], [40], [80], [30]]  # Single test score
y = ['pass', 'fail', 'pass', 'fail']
student_tree = tree.DecisionTreeClassifier()
student_tree = student_tree.fit(X, y)
print('Will a student with 60 pass?', student_tree.predict([[60]]))
Will a student with 60 pass? ['fail']

A Visual Look at Trees#

Decision trees are easier to understand with pictures.

Let us plot the student example tree.

# Let us plot the tree for the student example.
import matplotlib.pyplot as plt
plt.figure(figsize=(6,3))
tree.plot_tree(student_tree, feature_names=['score'], class_names=['fail','pass'], filled=True)
plt.show()
No description has been provided for this image
# What if there are more features for each student?
X2 = [[90, 1], [40, 0], [80, 1], [30, 0]]  # [score, extra_credit]
y2 = ['pass', 'fail', 'pass', 'fail']
tree2 = tree.DecisionTreeClassifier()
tree2.fit(X2, y2)
print('Student with score 60 and extra credit:', tree2.predict([[60, 1]]))
Student with score 60 and extra credit: ['fail']

Splitting Data for Training and Testing#

It is important to test the tree on new data to see if it learned well.

We should split our data into a training set and a testing set.

Let us do that now.

# Let us split our dataset.
from sklearn.model_selection import train_test_split
X = [[90], [40], [80], [30], [50], [70]]
y = ['pass', 'fail', 'pass', 'fail', 'fail', 'pass']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.33, random_state=42)
test_tree = tree.DecisionTreeClassifier()
test_tree.fit(X_train, y_train)
score = test_tree.score(X_test, y_test)
print('Test accuracy:', score)
Test accuracy: 1.0

Safe Access and Handling Missing Data#

Sometimes, your data has missing values.

You can use extra tools to fill gaps or skip missing data.

Let us see what happens when a value is missing.

# Suppose one value is missing: None means missing here.
X_missing = [[90], [None], [80], [30]]
y_missing = ['pass', 'fail', 'pass', 'fail']
try:
    tree_with_missing = tree.DecisionTreeClassifier()
    tree_with_missing.fit(X_missing, y_missing)
except ValueError as e:
    print('Error:', e)
    
# Let us fix missing values with SimpleImputer.
from sklearn.impute import SimpleImputer
imputer = SimpleImputer(strategy='mean')
X_fixed = imputer.fit_transform([[90], [None], [80], [30]])
tree_fixed = tree.DecisionTreeClassifier()
tree_fixed.fit(X_fixed, y_missing)
print('Fixed and trained on missing data!')
Fixed and trained on missing data!

Making Decisions: Questions at Each Split#

Each node in a decision tree asks a yes or no question.

Good questions split the data clearly.

Trees decide on the best questions by themselves during training.

Let us show how a tree splits using 'max_depth' to limit levels.#

deep_tree = tree.DecisionTreeClassifier(max_depth=1) deep_tree.fit([[80], [30], [55], [70]], ['pass', 'fail', 'fail', 'pass']) tree.plot_tree(deep_tree, feature_names=['score'], class_names=['fail','pass'], filled=True) plt.show()

# How can you use trees with more real-world data, like the famous iris flower dataset?
from sklearn.datasets import load_iris
iris = load_iris()
iris_tree = tree.DecisionTreeClassifier()
iris_tree.fit(iris.data, iris.target)
print('Predictions for first 5 iris flowers:', iris_tree.predict(iris.data[:5]))
Predictions for first 5 iris flowers: [0 0 0 0 0]
# What is a random forest? It is a group of several decision trees working together.
from sklearn.ensemble import RandomForestClassifier
rf = RandomForestClassifier(n_estimators=10, random_state=1)
rf.fit(X, y)
print('Random Forest prediction:', rf.predict([[60]]))
Random Forest prediction: ['fail']
# Let us compare tree and forest predictions on more examples.
test_scores = [[35], [65], [90]]
tree_preds = test_tree.predict(test_scores)
forest_preds = rf.predict(test_scores)
print('Tree predictions:', tree_preds)
print('Forest predictions:', forest_preds)
Tree predictions: ['fail' 'pass' 'pass']
Forest predictions: ['fail' 'pass' 'pass']
# Random forests are less likely to overfit.
from sklearn.metrics import accuracy_score
y_pred_tree = test_tree.predict(X_test)
y_pred_forest = rf.predict(X_test)
print('Tree test accuracy:', accuracy_score(y_test, y_pred_tree))
print('Forest test accuracy:', accuracy_score(y_test, y_pred_forest))
Tree test accuracy: 1.0
Forest test accuracy: 1.0
# Forests use randomness. You can control it using random_state for repeatable results.
my_forest = RandomForestClassifier(n_estimators=5, random_state=42)
my_forest.fit(X, y)
print('Consistent prediction:', my_forest.predict([[60]]))
Consistent prediction: ['fail']

Extra: Tuning Trees and Forests#

There are many ways to tune your models, like changing max_depth, n_estimators, or split rules.

Experiment with these to get better predictions for your data.

# Mini-project: Predicting Titanic Survival!
from sklearn.datasets import fetch_openml
titanic = fetch_openml('titanic', version=1, as_frame=True)
X = titanic.data[['age', 'fare']].fillna(0)
y = titanic.target
titanic_tree = tree.DecisionTreeClassifier(max_depth=3)
titanic_tree.fit(X, y)
print('Predicted survival for 10 random passengers:', titanic_tree.predict(X.head(10)))
Predicted survival for 10 random passengers: ['1' '1' '1' '1' '1' '0' '1' '0' '0' '0']
# You try it! Type your age and fare to see if the tree thinks you would survive.
your_age = int(input('Enter your age: '))
your_fare = float(input('Enter your fare: '))
print('Would you survive?', titanic_tree.predict([[your_age, your_fare]]))
Would you survive? ['0']
c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\sklearn\utils\validation.py:2749: UserWarning: X does not have valid feature names, but DecisionTreeClassifier was fitted with feature names
  warnings.warn(
 
# Mini-project challenge: Use a random forest on the Titanic data.
titanic_rf = RandomForestClassifier(n_estimators=50, max_depth=4, random_state=1)
titanic_rf.fit(X, y)
print('Random Forest survived prediction for 10 passengers:', titanic_rf.predict(X.head(10)))
Random Forest survived prediction for 10 passengers: ['1' '1' '1' '1' '1' '0' '1' '0' '0' '0']
# Challenge: What happens if you use only one feature?
X_age = titanic.data[['age']].fillna(0)
single_feature_tree = tree.DecisionTreeClassifier(max_depth=3)
single_feature_tree.fit(X_age, y)
print('Survival prediction with age only:', single_feature_tree.predict(X_age.head(5)))
Survival prediction with age only: ['0' '1' '1' '0' '0']
# What errors can happen, and how do we fix them?
try:
    tree.DecisionTreeClassifier().fit([['twenty']], ['pass'])
except ValueError as e:
    print('Error:', e)
    
Error: could not convert string to float: 'twenty'
# Quick tip: Use .feature_importances_ to see what your tree thinks matters most.
print('Feature importances (Titanic):', titanic_tree.feature_importances_)
Feature importances (Titanic): [0.05731102 0.94268898]

Recap and Next Steps#

You learned how computers use decision trees and random forests to make predictions.

You handled missing data, tested accuracy, and explored model tuning.

Try these tools on your own favorite dataset!

Like this lesson? Please like, share, or subscribe to our channel for more.

Comment below if you tried your own dataset!

Thanks for Learning with Python!#

Try a practice project using trees or forests on your hobby or work data.

And...

Do not forget to like, comment, subscribe, and share this video with friends!

See you in the next lesson!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.