Lesson 29 · Data Mining
Predicting Diabetes Using Ensemble Machine Learning Models: A Step-by-Step Guide
Ensemble learning helps us make better predictions by combining the strengths of multiple models. Today, let us learn the essentials of ensemble methods,…
- CourseData Mining
- Lesson29 of 31
- Video22 min
- FormatJupyter notebook · 12 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbIntroduction to Ensemble Learning: Week 89#
Ensemble learning helps us make better predictions by combining the strengths of multiple models.
Today, let us learn the essentials of ensemble methods, see examples with the Pima Indians Diabetes Dataset, and practice with easy, hands-on code.
This lesson will be very beginner-friendly. Let us get started!
# Data setup (Pima Indians Diabetes Dataset)
import warnings; warnings.filterwarnings("ignore")
import pandas as pd
import numpy as np
np.random.seed(42)
url = 'https://raw.githubusercontent.com/jbrownlee/Datasets/master/pima-indians-diabetes.data.csv'
cols = ['Pregnancies','Glucose','BloodPressure','SkinThickness','Insulin','BMI','DiabetesPedigree','Age','Outcome']
df = pd.read_csv(url, names=cols)
print(df.shape)
print(df.head(3))
Quick data check#
Let us take a moment to look at the dataset columns. Each row is a patient record, and the 'Outcome' column tells us if diabetes was diagnosed.
# Checking for missing values
print(df.isnull().sum())
# Basic stats: mean, median, and spread
print(df.describe())
What is Ensemble Learning?#
Ensemble learning is a technique where we combine several models to get more reliable predictions.
Imagine asking a group instead of just one person for advice. The groups answer is often more accurate!
Popular ensemble methods include bagging, boosting, and stacking.
# Preparing data for modeling
from sklearn.model_selection import train_test_split
X = df.drop('Outcome', axis=1)
y = df['Outcome']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, random_state=42)
Bagging: Bootstrap Aggregating#
Bagging stands for bootstrap aggregating.
It builds many simple models on slightly different samples of the data, then averages their answers. This helps reduce mistakes caused by chance or outliers.
A classic example is the Random Forest.
# Using a Decision Tree (single model)
from sklearn.tree import DecisionTreeClassifier
tree = DecisionTreeClassifier(random_state=42)
tree.fit(X_train, y_train)
print('Accuracy on test set:', tree.score(X_test, y_test))
# Using Random Forest for bagging
from sklearn.ensemble import RandomForestClassifier
rf = RandomForestClassifier(n_estimators=50, random_state=42)
rf.fit(X_train, y_train)
print('Accuracy on test set:', rf.score(X_test, y_test))
# Feature importance from Random Forest
import numpy as np
import matplotlib.pyplot as plt
importances = rf.feature_importances_
indices = np.argsort(importances)[::-1]
plt.title('Feature Importance (Random Forest)')
plt.bar(range(X.shape[1]), importances[indices], color='teal')
plt.xticks(range(X.shape[1]), X.columns[indices], rotation=45)
plt.tight_layout()
plt.show()
Boosting: Learning from Mistakes#
Boosting builds models one after another. Each new model pays extra attention to the hardest examplesthe ones previous models got wrong. Boosting can help even simple models perform much better.
A popular method here is AdaBoost or Gradient Boosting.
# Trying AdaBoost
from sklearn.ensemble import AdaBoostClassifier
ada = AdaBoostClassifier(n_estimators=50, random_state=42)
ada.fit(X_train, y_train)
print('AdaBoost test accuracy:', ada.score(X_test, y_test))
# Trying Gradient Boosting
from sklearn.ensemble import GradientBoostingClassifier
gb = GradientBoostingClassifier(n_estimators=50, random_state=42)
gb.fit(X_train, y_train)
print('Gradient Boosting test accuracy:', gb.score(X_test, y_test))
Stacking: Combining Different Models#
Stacking uses several different model types at oncelike trees, logistic regression, and k-nearest neighbors.
A second model then learns to blend their predictions into a single result.
Let us see a simple stacking example.
# Simple stacking ensemble
from sklearn.linear_model import LogisticRegression
from sklearn.neighbors import KNeighborsClassifier
from sklearn.ensemble import StackingClassifier
base_learners = [
('lr', LogisticRegression()),
('knn', KNeighborsClassifier()),
('dt', DecisionTreeClassifier(random_state=42))
]
stack = StackingClassifier(estimators=base_learners, final_estimator=LogisticRegression(), cv=5)
stack.fit(X_train, y_train)
print('Stacked model accuracy:', stack.score(X_test, y_test))
# Quick user experiment: try your own test size!
test_size_input = input("Choose a test set size between 0.2 and 0.4 (like 0.25): ")
test_size = float(test_size_input)
X_train2, X_test2, y_train2, y_test2 = train_test_split(X, y, test_size=test_size, random_state=42)
rf2 = RandomForestClassifier(n_estimators=50, random_state=42)
rf2.fit(X_train2, y_train2)
print('New test size random forest accuracy:', rf2.score(X_test2, y_test2))
Tips for ensemble learning#
- Ensembles often work best when the individual models are different.
- Too many models or too much complexity can cause overfittingwhere predictions work well on training data, but not on new data.
- Always test your ensemble on data the models did not see during training.
# Challenge: Can you improve accuracy even more?
from sklearn.model_selection import GridSearchCV
params = { 'n_estimators': [30, 50, 70], 'max_depth': [None, 3, 5] }
grid = GridSearchCV(RandomForestClassifier(random_state=42), params, cv=3)
grid.fit(X_train, y_train)
print('Best accuracy:', grid.best_score_)
print('Best parameters:', grid.best_params_)
Practice exercise#
Try swapping in AdaBoost or GradientBoostingClassifier for stackings final estimator, or add new base learners.
Which combination gives the best results for you?
Recap: What did we learn?#
- Ensemble learning combines models for stronger predictions.
- Bagging and boosting use different tricks to help accuracy.
- Stacking blends different model types.
Now you have built your first ensemble models!
Thanks for following along!#
Like this video and subscribe for more friendly, hands-on Python and data mining tutorials. See you in the next lesson!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



