Lesson 31 · Data Mining
Understanding AdaBoost and XGBoost: Key Boosting Techniques in Machine Learning
Welcome! This lesson explores two powerful boosting algorithms AdaBoost and XGBoost using the Pima Indians Diabetes dataset. Boosting combines several weak…
- CourseData Mining
- Lesson31 of 31
- Video17 min
- FormatJupyter notebook · 16 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbWeek 89: Boosting Techniques AdaBoost and XGBoost#
Welcome! This lesson explores two powerful boosting algorithms AdaBoost and XGBoost using the Pima Indians Diabetes dataset.
Boosting combines several weak learners to build a strong model. These methods are popular in data mining and machine learning.
We will start by loading real-world health data, prepare it, and finish by using boosting to predict diabetes.
Let us begin!
# Suppress warnings for cleaner outputs
import warnings; warnings.filterwarnings("ignore")
import numpy as np
np.random.seed(42)
Data setup#
We will use the Pima Indians Diabetes dataset. This famous medical dataset helps us predict diabetes based on health measures.
import pandas as pd
url = 'https://raw.githubusercontent.com/jbrownlee/Datasets/master/pima-indians-diabetes.data.csv'
cols = ['Pregnancies','Glucose','BloodPressure','SkinThickness','Insulin','BMI','DiabetesPedigree','Age','Outcome']
df = pd.read_csv(url, names=cols)
print(df.shape)
print(df.head(3))
Check for missing or suspicious values#
Zeroes in variables like glucose or BMI can mean missing data. Let us check how many zeros appear.
# Count zeros in key columns
for col in ['Glucose', 'BloodPressure', 'SkinThickness', 'Insulin', 'BMI']:
print(f"{col}: {(df[col] == 0).sum()} zeros")
# Replace zeros with NaN, then fill with median
import numpy as np
for col in ['Glucose', 'BloodPressure', 'SkinThickness', 'Insulin', 'BMI']:
df[col] = df[col].replace(0, np.nan)
df[col].fillna(df[col].median(), inplace=True)
print(df.isnull().sum())
Explore outcome balance#
Let us check how many patients have diabetes versus those who do not.
# Value counts for the target variable
print(df['Outcome'].value_counts())
# Quick plot of class balance
import matplotlib.pyplot as plt
df['Outcome'].value_counts().plot(kind='bar', color=['skyblue','orchid'])
plt.xlabel('Diabetes Present (1=Yes, 0=No)')
plt.title('Class Balance for Diabetes')
plt.show()
# Features and target split
X = df.drop('Outcome', axis=1)
y = df['Outcome']
# Train-test split
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
print(X_train.shape, X_test.shape)
What is Boosting?#
Boosting combines several simple models (weak learners) to build a smarter one.
Each learner focuses more on points that past models missed.
Boosting often improves prediction accuracy in real-world tasks like health or finance.
# AdaBoost setup
from sklearn.ensemble import AdaBoostClassifier
model_ada = AdaBoostClassifier(n_estimators=50, random_state=42)
model_ada.fit(X_train, y_train)
# AdaBoost prediction and accuracy
from sklearn.metrics import accuracy_score
y_pred_ada = model_ada.predict(X_test)
score_ada = accuracy_score(y_test, y_pred_ada)
print(f"AdaBoost Test Accuracy: {score_ada:.2f}")
What about XGBoost?#
XGBoost is an extremely popular and powerful boosting method.
It was used to win many data competitions.
# XGBoost setup
import xgboost as xgb
model_xgb = xgb.XGBClassifier(use_label_encoder=False, eval_metric='logloss', random_state=42)
model_xgb.fit(X_train, y_train)
# XGBoost prediction and accuracy
y_pred_xgb = model_xgb.predict(X_test)
score_xgb = accuracy_score(y_test, y_pred_xgb)
print(f"XGBoost Test Accuracy: {score_xgb:.2f}")
# Plot feature importances from XGBoost
import seaborn as sns
importances = model_xgb.feature_importances_
features = X.columns
sns.barplot(x=importances, y=features, palette='Blues_r')
plt.title('Feature Importance from XGBoost')
plt.show()
Which model won?#
Compare the AdaBoost and XGBoost accuracy scores from above.
Notice which model is better on this dataset and think about why.
# Confusion matrix for XGBoost
from sklearn.metrics import confusion_matrix, ConfusionMatrixDisplay
cm = confusion_matrix(y_test, y_pred_xgb)
disp = ConfusionMatrixDisplay(confusion_matrix=cm, display_labels=model_xgb.classes_)
disp.plot(cmap='Oranges')
plt.title('XGBoost Confusion Matrix')
plt.show()
# Ask the user to try predicting a patient
sample = X_test.iloc[0:1]
print('Patient data:')
print(sample)
user_input = input('Do you think this patient will have diabetes? Enter 1 for Yes, 0 for No: ')
xgb_pred = model_xgb.predict(sample)[0]
print(f'XGBoost predicts: {xgb_pred}')
print('Did you and the model agree?')
Tips for Boosting Success#
- Always check your data quality first.
- Avoid overfitting by limiting tree depth and tweaking estimators.
- Try more or fewer estimators for other datasets.
- Tuning learning rate can improve results.
Practice often with new datasets to build experience!
# Quick challenge: Can you tune XGBoost learning rate?
model_xgb2 = xgb.XGBClassifier(use_label_encoder=False, eval_metric='logloss', learning_rate=0.1, random_state=42)
model_xgb2.fit(X_train, y_train)
score_xgb2 = accuracy_score(y_test, model_xgb2.predict(X_test))
print(f"XGBoost (learning_rate=0.1) Accuracy: {score_xgb2:.2f}")
Recap!#
You loaded and cleaned a real medical dataset.
You ran both AdaBoost and XGBoost to predict diabetes.
You visualized features and model results.
Great job exploring boosting techniques in Python!
Thank you for joining this hands-on lesson!
Keep practicing by trying boosting on other datasets.
Subscribe on YouTube for more beginner-friendly, practical data mining guides!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



