Mathew K Analytics

Lesson 46 · Python Fundamentals

Understanding Decision Trees and Random Forests in Python for Bible Teaching Applications

In this lesson, we will discover how decision trees and random forests can help computers make choices based on data. We will use simple examples and a…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb
 

Welcome to Learning Decision Trees and Random Forests in Python!#

In this lesson, we will discover how decision trees and random forests can help computers make choices based on data.

We will use simple examples and a real-world dataset to help you understand.

Let us begin our journey into machine learning!

# Let us start by making sure warnings do not fill up our notebook
import warnings
warnings.filterwarnings("ignore")

What are Decision Trees?#

A decision tree is a machine learning method that sorts data by asking a series of questions.

It looks like a flowchart with branches.

You can use it for many things: finding if an email is spam, deciding if someone survived the Titanic, and more.

# Here is a basic Python variable assignment
tree = "A system for making yes or no decisions!"
print(tree)
A system for making yes or no decisions!

Let us look at a real dataset: The Titanic#

We will use data about Titanic passengers to predict survival.

First, let us download the data with pandas, a Python data library.

# Data setup
import pandas as pd
url = "https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv"
df = pd.read_csv(url)
print("Shape of data:", df.shape)
df.head()
Shape of data: (891, 12)
PassengerId Survived Pclass Name Sex Age SibSp Parch Ticket Fare Cabin Embarked
0 1 0 3 Braund, Mr. Owen Harris male 22.0 1 0 A/5 21171 7.2500 NaN S
1 2 1 1 Cumings, Mrs. John Bradley (Florence Briggs Th... female 38.0 1 0 PC 17599 71.2833 C85 C
2 3 1 3 Heikkinen, Miss. Laina female 26.0 0 0 STON/O2. 3101282 7.9250 NaN S
3 4 1 1 Futrelle, Mrs. Jacques Heath (Lily May Peel) female 35.0 1 0 113803 53.1000 C123 S
4 5 0 3 Allen, Mr. William Henry male 35.0 0 0 373450 8.0500 NaN S
# Let us see our target: who survived?
df["Survived"].value_counts()
Survived
0    549
1    342
Name: count, dtype: int64
# Let us split data into training and testing parts
from sklearn.model_selection import train_test_split
features = ["Pclass", "Sex", "Age", "Fare"]
# Convert text to numbers
df["Sex"] = df["Sex"].map({"male": 0, "female": 1})
df["Age"].fillna(df["Age"].median(), inplace=True)
X = df[features]
y = df["Survived"]
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
print("Train size:", X_train.shape, "Test size:", X_test.shape)
Train size: (712, 4) Test size: (179, 4)

Building Our First Decision Tree#

We will use scikit-learn to build a simple decision tree classifier.

# Build and train a decision tree classifier
from sklearn.tree import DecisionTreeClassifier
tree = DecisionTreeClassifier(random_state=42)
tree.fit(X_train, y_train)
print("Model trained!")
Model trained!
# Use our trained tree to predict
predictions = tree.predict(X_test)
print("Predictions:", predictions[:10])
Predictions: [0 1 1 1 1 0 1 0 0 1]
# How accurate is our tree?
from sklearn.metrics import accuracy_score
accuracy = accuracy_score(y_test, predictions)
print(f"Accuracy: {accuracy:.2f}")
Accuracy: 0.73

What are Random Forests?#

A random forest is a group of many decision trees.

Each tree gets a different piece of data and makes its own prediction.

The forest votes and the most common answer wins.

# Build and train a random forest
from sklearn.ensemble import RandomForestClassifier
forest = RandomForestClassifier(n_estimators=100, random_state=42)
forest.fit(X_train, y_train)
forest_preds = forest.predict(X_test)
print("Random forest predictions:", forest_preds[:10])
Random forest predictions: [0 0 1 1 0 1 1 0 1 1]
# Test how well the forest performs
forest_acc = accuracy_score(y_test, forest_preds)
print(f"Random forest accuracy: {forest_acc:.2f}")
Random forest accuracy: 0.79
# Which features mattered most?
import matplotlib.pyplot as plt
import numpy as np
feature_importances = forest.feature_importances_
features = np.array(features)
plt.bar(features, feature_importances)
plt.title("Feature Importance in Random Forest")
plt.ylabel("Importance")
plt.xlabel("Feature")
plt.show()
No description has been provided for this image
# Use input to try your own prediction!
pclass = int(input("Enter passenger class (1, 2, or 3): "))
sex = input("Enter sex (male or female): ")
age = float(input("Enter age: "))
fare = float(input("Enter fare: "))
sex_val = 0 if sex == "male" else 1
your_data = [[pclass, sex_val, age, fare]]
surv_pred = forest.predict(your_data)
if surv_pred[0]==1:
    print("Prediction: survived")
else:
    print("Prediction: did not survive")
    
Prediction: survived
 
# Challenge: Try out-of-bag score for your forest
forest_oob = RandomForestClassifier(n_estimators=100, oob_score=True, random_state=42)
forest_oob.fit(X_train, y_train)
print(f"Out-of-bag score: {forest_oob.oob_score_:.2f}")
Out-of-bag score: 0.80
# Common problems: missing values
import numpy as np
has_null = np.any(df[features].isnull())
print("Any missing values?", has_null)
Any missing values? False
# Handy: see more model details with classification report
from sklearn.metrics import classification_report
print(classification_report(y_test, forest_preds))
              precision    recall  f1-score   support

           0       0.81      0.84      0.83       105
           1       0.76      0.73      0.74        74

    accuracy                           0.79       179
   macro avg       0.79      0.78      0.79       179
weighted avg       0.79      0.79      0.79       179

# Best practice: set a random_state for reproducible results
model_a = DecisionTreeClassifier(random_state=123)
model_b = DecisionTreeClassifier(random_state=456)
model_a.fit(X_train, y_train)
model_b.fit(X_train, y_train)
pred_a = model_a.predict(X_test)
pred_b = model_b.predict(X_test)
print(pred_a[:5], pred_b[:5])
[0 1 1 1 1] [0 1 1 1 1]

Challenge: Can you change the model to use a different feature?#

Try swapping out a feature or adding one more column.

What happens to your model's predictions?

Recap#

You learned to build decision trees and random forests.

You used them to predict Titanic survival.

Changing features and model settings helped you improve results.

You saw how data shape, accuracy, and feature importance work together.

Keep practicing and try your own datasets!

Thanks for learning with us!#

If you enjoyed this, please like, comment, and subscribe to our YouTube channel for more beginner-friendly Python tutorials.

Happy coding!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.