Mathew K Analytics

Lesson 31 · Data Mining

Understanding AdaBoost and XGBoost: Key Boosting Techniques in Machine Learning

Welcome! This lesson explores two powerful boosting algorithms AdaBoost and XGBoost using the Pima Indians Diabetes dataset. Boosting combines several weak…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Week 89: Boosting Techniques AdaBoost and XGBoost#

Welcome! This lesson explores two powerful boosting algorithms AdaBoost and XGBoost using the Pima Indians Diabetes dataset.

Boosting combines several weak learners to build a strong model. These methods are popular in data mining and machine learning.

We will start by loading real-world health data, prepare it, and finish by using boosting to predict diabetes.

Let us begin!

# Suppress warnings for cleaner outputs
import warnings; warnings.filterwarnings("ignore")
import numpy as np
np.random.seed(42)

Data setup#

We will use the Pima Indians Diabetes dataset. This famous medical dataset helps us predict diabetes based on health measures.

import pandas as pd
url = 'https://raw.githubusercontent.com/jbrownlee/Datasets/master/pima-indians-diabetes.data.csv'
cols = ['Pregnancies','Glucose','BloodPressure','SkinThickness','Insulin','BMI','DiabetesPedigree','Age','Outcome']
df = pd.read_csv(url, names=cols)
print(df.shape)
print(df.head(3))
(768, 9)
   Pregnancies  Glucose  BloodPressure  SkinThickness  Insulin   BMI  \
0            6      148             72             35        0  33.6   
1            1       85             66             29        0  26.6   
2            8      183             64              0        0  23.3   

   DiabetesPedigree  Age  Outcome  
0             0.627   50        1  
1             0.351   31        0  
2             0.672   32        1  

Check for missing or suspicious values#

Zeroes in variables like glucose or BMI can mean missing data. Let us check how many zeros appear.

# Count zeros in key columns
for col in ['Glucose', 'BloodPressure', 'SkinThickness', 'Insulin', 'BMI']:
    print(f"{col}: {(df[col] == 0).sum()} zeros")
    
Glucose: 5 zeros
BloodPressure: 35 zeros
SkinThickness: 227 zeros
Insulin: 374 zeros
BMI: 11 zeros
# Replace zeros with NaN, then fill with median
import numpy as np
for col in ['Glucose', 'BloodPressure', 'SkinThickness', 'Insulin', 'BMI']:
    df[col] = df[col].replace(0, np.nan)
    df[col].fillna(df[col].median(), inplace=True)
print(df.isnull().sum())
Pregnancies         0
Glucose             0
BloodPressure       0
SkinThickness       0
Insulin             0
BMI                 0
DiabetesPedigree    0
Age                 0
Outcome             0
dtype: int64

Explore outcome balance#

Let us check how many patients have diabetes versus those who do not.

# Value counts for the target variable
print(df['Outcome'].value_counts())
Outcome
0    500
1    268
Name: count, dtype: int64
# Quick plot of class balance
import matplotlib.pyplot as plt
df['Outcome'].value_counts().plot(kind='bar', color=['skyblue','orchid'])
plt.xlabel('Diabetes Present (1=Yes, 0=No)')
plt.title('Class Balance for Diabetes')
plt.show()
No description has been provided for this image
# Features and target split
X = df.drop('Outcome', axis=1)
y = df['Outcome']
# Train-test split
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
print(X_train.shape, X_test.shape)
(614, 8) (154, 8)

What is Boosting?#

Boosting combines several simple models (weak learners) to build a smarter one.

Each learner focuses more on points that past models missed.

Boosting often improves prediction accuracy in real-world tasks like health or finance.

# AdaBoost setup
from sklearn.ensemble import AdaBoostClassifier
model_ada = AdaBoostClassifier(n_estimators=50, random_state=42)
model_ada.fit(X_train, y_train)
AdaBoostClassifier(random_state=42)
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
# AdaBoost prediction and accuracy
from sklearn.metrics import accuracy_score
y_pred_ada = model_ada.predict(X_test)
score_ada = accuracy_score(y_test, y_pred_ada)
print(f"AdaBoost Test Accuracy: {score_ada:.2f}")
AdaBoost Test Accuracy: 0.75

What about XGBoost?#

XGBoost is an extremely popular and powerful boosting method.

It was used to win many data competitions.

# XGBoost setup
import xgboost as xgb
model_xgb = xgb.XGBClassifier(use_label_encoder=False, eval_metric='logloss', random_state=42)
model_xgb.fit(X_train, y_train)
XGBClassifier(base_score=None, booster=None, callbacks=None,
              colsample_bylevel=None, colsample_bynode=None,
              colsample_bytree=None, device=None, early_stopping_rounds=None,
              enable_categorical=False, eval_metric='logloss',
              feature_types=None, feature_weights=None, gamma=None,
              grow_policy=None, importance_type=None,
              interaction_constraints=None, learning_rate=None, max_bin=None,
              max_cat_threshold=None, max_cat_to_onehot=None,
              max_delta_step=None, max_depth=None, max_leaves=None,
              min_child_weight=None, missing=nan, monotone_constraints=None,
              multi_strategy=None, n_estimators=None, n_jobs=None,
              num_parallel_tree=None, ...)
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
# XGBoost prediction and accuracy
y_pred_xgb = model_xgb.predict(X_test)
score_xgb = accuracy_score(y_test, y_pred_xgb)
print(f"XGBoost Test Accuracy: {score_xgb:.2f}")
XGBoost Test Accuracy: 0.71
# Plot feature importances from XGBoost
import seaborn as sns
importances = model_xgb.feature_importances_
features = X.columns
sns.barplot(x=importances, y=features, palette='Blues_r')
plt.title('Feature Importance from XGBoost')
plt.show()
No description has been provided for this image

Which model won?#

Compare the AdaBoost and XGBoost accuracy scores from above.

Notice which model is better on this dataset and think about why.

# Confusion matrix for XGBoost
from sklearn.metrics import confusion_matrix, ConfusionMatrixDisplay
cm = confusion_matrix(y_test, y_pred_xgb)
disp = ConfusionMatrixDisplay(confusion_matrix=cm, display_labels=model_xgb.classes_)
disp.plot(cmap='Oranges')
plt.title('XGBoost Confusion Matrix')
plt.show()
No description has been provided for this image
# Ask the user to try predicting a patient
sample = X_test.iloc[0:1]
print('Patient data:')
print(sample)
user_input = input('Do you think this patient will have diabetes? Enter 1 for Yes, 0 for No: ')
xgb_pred = model_xgb.predict(sample)[0]
print(f'XGBoost predicts: {xgb_pred}')
print('Did you and the model agree?')
Patient data:
     Pregnancies  Glucose  BloodPressure  SkinThickness  Insulin   BMI  \
668            6     98.0           58.0           33.0    190.0  34.0   

     DiabetesPedigree  Age  
668              0.43   43  
XGBoost predicts: 1
Did you and the model agree?

Tips for Boosting Success#

  1. Always check your data quality first.
  2. Avoid overfitting by limiting tree depth and tweaking estimators.
  3. Try more or fewer estimators for other datasets.
  4. Tuning learning rate can improve results.

Practice often with new datasets to build experience!

# Quick challenge: Can you tune XGBoost learning rate?
model_xgb2 = xgb.XGBClassifier(use_label_encoder=False, eval_metric='logloss', learning_rate=0.1, random_state=42)
model_xgb2.fit(X_train, y_train)
score_xgb2 = accuracy_score(y_test, model_xgb2.predict(X_test))
print(f"XGBoost (learning_rate=0.1) Accuracy: {score_xgb2:.2f}")
XGBoost (learning_rate=0.1) Accuracy: 0.72

Recap!#

You loaded and cleaned a real medical dataset.

You ran both AdaBoost and XGBoost to predict diabetes.

You visualized features and model results.

Great job exploring boosting techniques in Python!

Thank you for joining this hands-on lesson!

Keep practicing by trying boosting on other datasets.

Subscribe on YouTube for more beginner-friendly, practical data mining guides!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.