Mathew K Analytics

Lesson 29 · Data Mining

Predicting Diabetes Using Ensemble Machine Learning Models: A Step-by-Step Guide

Ensemble learning helps us make better predictions by combining the strengths of multiple models. Today, let us learn the essentials of ensemble methods,…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Introduction to Ensemble Learning: Week 89#

Ensemble learning helps us make better predictions by combining the strengths of multiple models.

Today, let us learn the essentials of ensemble methods, see examples with the Pima Indians Diabetes Dataset, and practice with easy, hands-on code.

This lesson will be very beginner-friendly. Let us get started!

# Data setup (Pima Indians Diabetes Dataset)
import warnings; warnings.filterwarnings("ignore")
import pandas as pd
import numpy as np
np.random.seed(42)
url = 'https://raw.githubusercontent.com/jbrownlee/Datasets/master/pima-indians-diabetes.data.csv'
cols = ['Pregnancies','Glucose','BloodPressure','SkinThickness','Insulin','BMI','DiabetesPedigree','Age','Outcome']
df = pd.read_csv(url, names=cols)
print(df.shape)
print(df.head(3))
(768, 9)
   Pregnancies  Glucose  BloodPressure  SkinThickness  Insulin   BMI  \
0            6      148             72             35        0  33.6   
1            1       85             66             29        0  26.6   
2            8      183             64              0        0  23.3   

   DiabetesPedigree  Age  Outcome  
0             0.627   50        1  
1             0.351   31        0  
2             0.672   32        1  

Quick data check#

Let us take a moment to look at the dataset columns. Each row is a patient record, and the 'Outcome' column tells us if diabetes was diagnosed.

# Checking for missing values
print(df.isnull().sum())
Pregnancies         0
Glucose             0
BloodPressure       0
SkinThickness       0
Insulin             0
BMI                 0
DiabetesPedigree    0
Age                 0
Outcome             0
dtype: int64
# Basic stats: mean, median, and spread
print(df.describe())
       Pregnancies     Glucose  BloodPressure  SkinThickness     Insulin  \
count   768.000000  768.000000     768.000000     768.000000  768.000000   
mean      3.845052  120.894531      69.105469      20.536458   79.799479   
std       3.369578   31.972618      19.355807      15.952218  115.244002   
min       0.000000    0.000000       0.000000       0.000000    0.000000   
25%       1.000000   99.000000      62.000000       0.000000    0.000000   
50%       3.000000  117.000000      72.000000      23.000000   30.500000   
75%       6.000000  140.250000      80.000000      32.000000  127.250000   
max      17.000000  199.000000     122.000000      99.000000  846.000000   

              BMI  DiabetesPedigree         Age     Outcome  
count  768.000000        768.000000  768.000000  768.000000  
mean    31.992578          0.471876   33.240885    0.348958  
std      7.884160          0.331329   11.760232    0.476951  
min      0.000000          0.078000   21.000000    0.000000  
25%     27.300000          0.243750   24.000000    0.000000  
50%     32.000000          0.372500   29.000000    0.000000  
75%     36.600000          0.626250   41.000000    1.000000  
max     67.100000          2.420000   81.000000    1.000000  

What is Ensemble Learning?#

Ensemble learning is a technique where we combine several models to get more reliable predictions.

Imagine asking a group instead of just one person for advice. The groups answer is often more accurate!

Popular ensemble methods include bagging, boosting, and stacking.

# Preparing data for modeling
from sklearn.model_selection import train_test_split
X = df.drop('Outcome', axis=1)
y = df['Outcome']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, random_state=42)

Bagging: Bootstrap Aggregating#

Bagging stands for bootstrap aggregating.

It builds many simple models on slightly different samples of the data, then averages their answers. This helps reduce mistakes caused by chance or outliers.

A classic example is the Random Forest.

# Using a Decision Tree (single model)
from sklearn.tree import DecisionTreeClassifier
tree = DecisionTreeClassifier(random_state=42)
tree.fit(X_train, y_train)
print('Accuracy on test set:', tree.score(X_test, y_test))
Accuracy on test set: 0.7083333333333334
# Using Random Forest for bagging
from sklearn.ensemble import RandomForestClassifier
rf = RandomForestClassifier(n_estimators=50, random_state=42)
rf.fit(X_train, y_train)
print('Accuracy on test set:', rf.score(X_test, y_test))
Accuracy on test set: 0.7395833333333334
# Feature importance from Random Forest
import numpy as np
import matplotlib.pyplot as plt
importances = rf.feature_importances_
indices = np.argsort(importances)[::-1]
plt.title('Feature Importance (Random Forest)')
plt.bar(range(X.shape[1]), importances[indices], color='teal')
plt.xticks(range(X.shape[1]), X.columns[indices], rotation=45)
plt.tight_layout()
plt.show()
No description has been provided for this image

Boosting: Learning from Mistakes#

Boosting builds models one after another. Each new model pays extra attention to the hardest examplesthe ones previous models got wrong. Boosting can help even simple models perform much better.

A popular method here is AdaBoost or Gradient Boosting.

# Trying AdaBoost
from sklearn.ensemble import AdaBoostClassifier
ada = AdaBoostClassifier(n_estimators=50, random_state=42)
ada.fit(X_train, y_train)
print('AdaBoost test accuracy:', ada.score(X_test, y_test))
AdaBoost test accuracy: 0.734375
# Trying Gradient Boosting
from sklearn.ensemble import GradientBoostingClassifier
gb = GradientBoostingClassifier(n_estimators=50, random_state=42)
gb.fit(X_train, y_train)
print('Gradient Boosting test accuracy:', gb.score(X_test, y_test))
Gradient Boosting test accuracy: 0.75

Stacking: Combining Different Models#

Stacking uses several different model types at oncelike trees, logistic regression, and k-nearest neighbors.

A second model then learns to blend their predictions into a single result.

Let us see a simple stacking example.

# Simple stacking ensemble
from sklearn.linear_model import LogisticRegression
from sklearn.neighbors import KNeighborsClassifier
from sklearn.ensemble import StackingClassifier
base_learners = [
    ('lr', LogisticRegression()),
    ('knn', KNeighborsClassifier()),
    ('dt', DecisionTreeClassifier(random_state=42))
]
stack = StackingClassifier(estimators=base_learners, final_estimator=LogisticRegression(), cv=5)
stack.fit(X_train, y_train)
print('Stacked model accuracy:', stack.score(X_test, y_test))
Stacked model accuracy: 0.7447916666666666
# Quick user experiment: try your own test size!
test_size_input = input("Choose a test set size between 0.2 and 0.4 (like 0.25): ")
test_size = float(test_size_input)
X_train2, X_test2, y_train2, y_test2 = train_test_split(X, y, test_size=test_size, random_state=42)
rf2 = RandomForestClassifier(n_estimators=50, random_state=42)
rf2.fit(X_train2, y_train2)
print('New test size random forest accuracy:', rf2.score(X_test2, y_test2))
New test size random forest accuracy: 0.7402597402597403

Tips for ensemble learning#

  • Ensembles often work best when the individual models are different.
  • Too many models or too much complexity can cause overfittingwhere predictions work well on training data, but not on new data.
  • Always test your ensemble on data the models did not see during training.
# Challenge: Can you improve accuracy even more?
from sklearn.model_selection import GridSearchCV
params = { 'n_estimators': [30, 50, 70], 'max_depth': [None, 3, 5] }
grid = GridSearchCV(RandomForestClassifier(random_state=42), params, cv=3)
grid.fit(X_train, y_train)
print('Best accuracy:', grid.best_score_)
print('Best parameters:', grid.best_params_)
Best accuracy: 0.7690972222222222
Best parameters: {'max_depth': 5, 'n_estimators': 30}

Practice exercise#

Try swapping in AdaBoost or GradientBoostingClassifier for stackings final estimator, or add new base learners.

Which combination gives the best results for you?

Recap: What did we learn?#

  • Ensemble learning combines models for stronger predictions.
  • Bagging and boosting use different tricks to help accuracy.
  • Stacking blends different model types.

Now you have built your first ensemble models!

Thanks for following along!#

Like this video and subscribe for more friendly, hands-on Python and data mining tutorials. See you in the next lesson!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.