Lesson 30 · Data Mining
Understanding Bagging and Random Forests: Ensemble Methods in Machine Learning
Welcome! Today we will learn about ensemble methods: bagging and random forests. These methods help improve prediction accuracy by combining multiple…
- CourseData Mining
- Lesson30 of 31
- Video23 min
- FormatJupyter notebook · 18 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbWeek 89: Introduction to Bagging and Random Forests#
Welcome! Today we will learn about ensemble methods: bagging and random forests. These methods help improve prediction accuracy by combining multiple models.
We will use real diabetes data and apply bagging and random forests to predict diabetes risk.
Let us begin our data mining journey!
import warnings; warnings.filterwarnings("ignore")
import numpy as np
np.random.seed(42)
# This will suppress warnings to keep our notebook clean
# Data setup (Pima Indians Diabetes Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/jbrownlee/Datasets/master/pima-indians-diabetes.data.csv'
cols = ['Pregnancies','Glucose','BloodPressure','SkinThickness','Insulin','BMI','DiabetesPedigree','Age','Outcome']
df = pd.read_csv(url, names=cols)
print(df.shape)
print(df.head(3))
# Checking for missing values
df.isnull().sum()
# Checking for strange zeros (as missing) in some columns
cols_with_zeros = ['Glucose', 'BloodPressure', 'SkinThickness', 'Insulin', 'BMI']
for col in cols_with_zeros:
print(f"{col}:", (df[col] == 0).sum())
# Replacing zeros with NaN in suspicious columns
import numpy as np
for col in cols_with_zeros:
df[col] = df[col].replace(0, np.nan)
df[cols_with_zeros].isnull().sum()
# Filling missing values with the column median
df.fillna(df.median(), inplace=True)
# Summary statistics
df.describe()
# Plot histograms for each feature
import matplotlib.pyplot as plt
df.hist(figsize=(10,7)); plt.tight_layout(); plt.show()
Bagging: Bootstrap Aggregating Explained#
Bagging builds many models from random pieces of the data. Votes from these models are combined for the final answer.
This can stabilize and boost performance, especially for unstable models.
We will try bagging with decision trees next!
# Split the data into training and testing sets
from sklearn.model_selection import train_test_split
X = df.drop('Outcome', axis=1)
y = df['Outcome']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)
print(X_train.shape, X_test.shape)
# Bagging Classifier with Decision Trees
from sklearn.ensemble import BaggingClassifier
from sklearn.tree import DecisionTreeClassifier
bag = BaggingClassifier(DecisionTreeClassifier(), n_estimators=50, random_state=42)
bag.fit(X_train, y_train)
bag_score = bag.score(X_test, y_test)
print(f"Bagging Test Accuracy: {bag_score:.3f}")
# Compare: A single Decision Tree
tree = DecisionTreeClassifier(random_state=42)
tree.fit(X_train, y_train)
tree_score = tree.score(X_test, y_test)
print(f"Single Tree Test Accuracy: {tree_score:.3f}")
Random Forests: Smart Bagging#
Random Forests are a form of bagging, but smarter. They add randomness by letting each tree look at only some features at each split.
This helps make the trees more different from each other.
It is a very popular and powerful method for both classification and regression.
# Random Forest Classifier
from sklearn.ensemble import RandomForestClassifier
rf = RandomForestClassifier(n_estimators=100, random_state=42)
rf.fit(X_train, y_train)
rf_score = rf.score(X_test, y_test)
print(f"Random Forest Test Accuracy: {rf_score:.3f}")
# What features mattered most? Feature importance
importances = rf.feature_importances_
for col, imp in zip(X.columns, importances):
print(f"{col}: {imp:.3f}")
# Visual confusion matrix for Random Forest
from sklearn.metrics import confusion_matrix, ConfusionMatrixDisplay
y_pred = rf.predict(X_test)
cm = confusion_matrix(y_test, y_pred)
disp = ConfusionMatrixDisplay(cm)
disp.plot(); plt.show()
# Interactive: Make your own diabetes prediction
data_point = []
for col in X.columns:
value = input(f"Enter value for {col}: ")
data_point.append(float(value))
prediction = rf.predict([data_point])[0]
result = 'Diabetes' if prediction == 1 else 'No Diabetes'
print(f"Prediction: {result}")
Mini-project: Bagging versus Random Forests#
Let us see if bagging and random forests predict well for a new set of cases. You can try different numbers of trees or change the feature set to see what happens.
- Could you beat our best score with a new strategy?
- What might you change next?
Write your ideas below or try with the code!
# Challenge: Bagging with more or fewer trees
for n in [10, 50, 200]:
bag = BaggingClassifier(DecisionTreeClassifier(), n_estimators=n, random_state=42)
bag.fit(X_train, y_train)
bag_score = bag.score(X_test, y_test)
print(f"Bagging with {n} trees, Test Accuracy: {bag_score:.3f}")
# Extra tip: Save and reload your Random Forest
import joblib
joblib.dump(rf, 'rf_model.joblib')
rf_new = joblib.load('rf_model.joblib')
print(f"Reloaded forest, Test Accuracy: {rf_new.score(X_test, y_test):.3f}")
Troubleshooting and Best Practices#
When to use bagging or random forests:
- When single models overfit your data.
- When accuracy is unstable across runs.
Always check the importance of your features. Use random_state for reproducibility. Try tuning the number of trees or tree depth.
If accuracy is low, check for data quality or biased splits.
Ask for help on forums or from teachers if you are stuck!
# Practice: Try another classifier (optional)
from sklearn.neighbors import KNeighborsClassifier
knn = KNeighborsClassifier()
knn.fit(X_train, y_train)
knn_score = knn.score(X_test, y_test)
print(f"kNN Test Accuracy: {knn_score:.3f}")
Recap: What We Learned#
- Bagging helps average out errors with many trees.
- Random forests add feature randomness, making trees more diverse.
- These methods are reliable for tricky or noisy data.
- Checking feature importance can give you insights for real decisions.
- Always compare ensemble methods with single models.
Congratulations! You are building skills found in industry data science jobs.
Thank You and Next Steps#
If you learned something new, hit thumbs up and subscribe to the channel!
Try the challenge exercises for extra practice. Ask questions in the comments and share your own results.
Happy mining!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



