Mathew K Analytics

Lesson 30 · Data Mining

Understanding Bagging and Random Forests: Ensemble Methods in Machine Learning

Welcome! Today we will learn about ensemble methods: bagging and random forests. These methods help improve prediction accuracy by combining multiple…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Week 89: Introduction to Bagging and Random Forests#

Welcome! Today we will learn about ensemble methods: bagging and random forests. These methods help improve prediction accuracy by combining multiple models.

We will use real diabetes data and apply bagging and random forests to predict diabetes risk.

Let us begin our data mining journey!

import warnings; warnings.filterwarnings("ignore")
import numpy as np
np.random.seed(42)
# This will suppress warnings to keep our notebook clean
# Data setup (Pima Indians Diabetes Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/jbrownlee/Datasets/master/pima-indians-diabetes.data.csv'
cols = ['Pregnancies','Glucose','BloodPressure','SkinThickness','Insulin','BMI','DiabetesPedigree','Age','Outcome']
df = pd.read_csv(url, names=cols)
print(df.shape)
print(df.head(3))
(768, 9)
   Pregnancies  Glucose  BloodPressure  SkinThickness  Insulin   BMI  \
0            6      148             72             35        0  33.6   
1            1       85             66             29        0  26.6   
2            8      183             64              0        0  23.3   

   DiabetesPedigree  Age  Outcome  
0             0.627   50        1  
1             0.351   31        0  
2             0.672   32        1  
# Checking for missing values
df.isnull().sum()
Pregnancies         0
Glucose             0
BloodPressure       0
SkinThickness       0
Insulin             0
BMI                 0
DiabetesPedigree    0
Age                 0
Outcome             0
dtype: int64
# Checking for strange zeros (as missing) in some columns
cols_with_zeros = ['Glucose', 'BloodPressure', 'SkinThickness', 'Insulin', 'BMI']
for col in cols_with_zeros:
    print(f"{col}:", (df[col] == 0).sum())
    
Glucose: 5
BloodPressure: 35
SkinThickness: 227
Insulin: 374
BMI: 11
# Replacing zeros with NaN in suspicious columns
import numpy as np
for col in cols_with_zeros:
    df[col] = df[col].replace(0, np.nan)
df[cols_with_zeros].isnull().sum()
Glucose            5
BloodPressure     35
SkinThickness    227
Insulin          374
BMI               11
dtype: int64
# Filling missing values with the column median
df.fillna(df.median(), inplace=True)
# Summary statistics
df.describe()
Pregnancies Glucose BloodPressure SkinThickness Insulin BMI DiabetesPedigree Age Outcome
count 768.000000 768.000000 768.000000 768.000000 768.000000 768.000000 768.000000 768.000000 768.000000
mean 3.845052 121.656250 72.386719 29.108073 140.671875 32.455208 0.471876 33.240885 0.348958
std 3.369578 30.438286 12.096642 8.791221 86.383060 6.875177 0.331329 11.760232 0.476951
min 0.000000 44.000000 24.000000 7.000000 14.000000 18.200000 0.078000 21.000000 0.000000
25% 1.000000 99.750000 64.000000 25.000000 121.500000 27.500000 0.243750 24.000000 0.000000
50% 3.000000 117.000000 72.000000 29.000000 125.000000 32.300000 0.372500 29.000000 0.000000
75% 6.000000 140.250000 80.000000 32.000000 127.250000 36.600000 0.626250 41.000000 1.000000
max 17.000000 199.000000 122.000000 99.000000 846.000000 67.100000 2.420000 81.000000 1.000000
# Plot histograms for each feature
import matplotlib.pyplot as plt
df.hist(figsize=(10,7)); plt.tight_layout(); plt.show()
No description has been provided for this image

Bagging: Bootstrap Aggregating Explained#

Bagging builds many models from random pieces of the data. Votes from these models are combined for the final answer.

This can stabilize and boost performance, especially for unstable models.

We will try bagging with decision trees next!

# Split the data into training and testing sets
from sklearn.model_selection import train_test_split
X = df.drop('Outcome', axis=1)
y = df['Outcome']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)
print(X_train.shape, X_test.shape)
(537, 8) (231, 8)
# Bagging Classifier with Decision Trees
from sklearn.ensemble import BaggingClassifier
from sklearn.tree import DecisionTreeClassifier
bag = BaggingClassifier(DecisionTreeClassifier(), n_estimators=50, random_state=42)
bag.fit(X_train, y_train)
bag_score = bag.score(X_test, y_test)
print(f"Bagging Test Accuracy: {bag_score:.3f}")
Bagging Test Accuracy: 0.740
# Compare: A single Decision Tree
tree = DecisionTreeClassifier(random_state=42)
tree.fit(X_train, y_train)
tree_score = tree.score(X_test, y_test)
print(f"Single Tree Test Accuracy: {tree_score:.3f}")
Single Tree Test Accuracy: 0.697

Random Forests: Smart Bagging#

Random Forests are a form of bagging, but smarter. They add randomness by letting each tree look at only some features at each split.

This helps make the trees more different from each other.

It is a very popular and powerful method for both classification and regression.

# Random Forest Classifier
from sklearn.ensemble import RandomForestClassifier
rf = RandomForestClassifier(n_estimators=100, random_state=42)
rf.fit(X_train, y_train)
rf_score = rf.score(X_test, y_test)
print(f"Random Forest Test Accuracy: {rf_score:.3f}")
Random Forest Test Accuracy: 0.749
# What features mattered most? Feature importance
importances = rf.feature_importances_
for col, imp in zip(X.columns, importances):
    print(f"{col}: {imp:.3f}")
    
Pregnancies: 0.079
Glucose: 0.283
BloodPressure: 0.082
SkinThickness: 0.070
Insulin: 0.080
BMI: 0.156
DiabetesPedigree: 0.111
Age: 0.140
# Visual confusion matrix for Random Forest
from sklearn.metrics import confusion_matrix, ConfusionMatrixDisplay
y_pred = rf.predict(X_test)
cm = confusion_matrix(y_test, y_pred)
disp = ConfusionMatrixDisplay(cm)
disp.plot(); plt.show()
No description has been provided for this image
# Interactive: Make your own diabetes prediction
data_point = []
for col in X.columns:
    value = input(f"Enter value for {col}: ")
    data_point.append(float(value))
prediction = rf.predict([data_point])[0]
result = 'Diabetes' if prediction == 1 else 'No Diabetes'
print(f"Prediction: {result}")
Prediction: No Diabetes

Mini-project: Bagging versus Random Forests#

Let us see if bagging and random forests predict well for a new set of cases. You can try different numbers of trees or change the feature set to see what happens.

  • Could you beat our best score with a new strategy?
  • What might you change next?

Write your ideas below or try with the code!

# Challenge: Bagging with more or fewer trees
for n in [10, 50, 200]:
    bag = BaggingClassifier(DecisionTreeClassifier(), n_estimators=n, random_state=42)
    bag.fit(X_train, y_train)
    bag_score = bag.score(X_test, y_test)
    print(f"Bagging with {n} trees, Test Accuracy: {bag_score:.3f}")
    
Bagging with 10 trees, Test Accuracy: 0.762
Bagging with 50 trees, Test Accuracy: 0.740
Bagging with 200 trees, Test Accuracy: 0.753
# Extra tip: Save and reload your Random Forest
import joblib
joblib.dump(rf, 'rf_model.joblib')
rf_new = joblib.load('rf_model.joblib')
print(f"Reloaded forest, Test Accuracy: {rf_new.score(X_test, y_test):.3f}")
Reloaded forest, Test Accuracy: 0.749

Troubleshooting and Best Practices#

When to use bagging or random forests:

  • When single models overfit your data.
  • When accuracy is unstable across runs.

Always check the importance of your features. Use random_state for reproducibility. Try tuning the number of trees or tree depth.

If accuracy is low, check for data quality or biased splits.

Ask for help on forums or from teachers if you are stuck!

# Practice: Try another classifier (optional)
from sklearn.neighbors import KNeighborsClassifier
knn = KNeighborsClassifier()
knn.fit(X_train, y_train)
knn_score = knn.score(X_test, y_test)
print(f"kNN Test Accuracy: {knn_score:.3f}")
kNN Test Accuracy: 0.675

Recap: What We Learned#

  • Bagging helps average out errors with many trees.
  • Random forests add feature randomness, making trees more diverse.
  • These methods are reliable for tricky or noisy data.
  • Checking feature importance can give you insights for real decisions.
  • Always compare ensemble methods with single models.

Congratulations! You are building skills found in industry data science jobs.

Thank You and Next Steps#

If you learned something new, hit thumbs up and subscribe to the channel!

Try the challenge exercises for extra practice. Ask questions in the comments and share your own results.

Happy mining!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.