Mathew K Analytics

Lesson 75 · Python Fundamentals

Understanding Gradient Boosting and Implementing XGBoost in Python for Predictive Modeling

In this lesson, you will learn the basics of how to use gradient boosting with the powerful XGBoost library. We will work step-by-step, using a real-world…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb
 

Welcome to Gradient Boosting with XGBoost in Python!#

In this lesson, you will learn the basics of how to use gradient boosting with the powerful XGBoost library.

We will work step-by-step, using a real-world dataset to predict who survived the Titanic.

You do not need to know machine learning yet!

Let us get started and see how simple it can be.

# First, let us import the packages we will need.
import warnings
warnings.filterwarnings('ignore')
import pandas as pd
import numpy as np
import xgboost as xgb
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score
---------------------------------------------------------------------------
ModuleNotFoundError                       Traceback (most recent call last)
Cell In[1], line 6
      4 import pandas as pd
      5 import numpy as np
----> 6 import xgboost as xgb
      7 from sklearn.model_selection import train_test_split
      8 from sklearn.metrics import accuracy_score

ModuleNotFoundError: No module named 'xgboost'

What is Gradient Boosting?#

Gradient boosting is a clever way to make strong predictions by combining many simple models.

Each simple model learns from the mistakes of the last one.

This helps the final predictions get better and better.

# Data setup: Let us load the Titanic dataset from the web.
url = "https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv"
data = pd.read_csv(url)
print("Shape:", data.shape)
data.head()
Shape: (891, 12)
PassengerId Survived Pclass Name Sex Age SibSp Parch Ticket Fare Cabin Embarked
0 1 0 3 Braund, Mr. Owen Harris male 22.0 1 0 A/5 21171 7.2500 NaN S
1 2 1 1 Cumings, Mrs. John Bradley (Florence Briggs Th... female 38.0 1 0 PC 17599 71.2833 C85 C
2 3 1 3 Heikkinen, Miss. Laina female 26.0 0 0 STON/O2. 3101282 7.9250 NaN S
3 4 1 1 Futrelle, Mrs. Jacques Heath (Lily May Peel) female 35.0 1 0 113803 53.1000 C123 S
4 5 0 3 Allen, Mr. William Henry male 35.0 0 0 373450 8.0500 NaN S
# Let us check how many people survived and how many did not.
data['Survived'].value_counts()
Survived
0    549
1    342
Name: count, dtype: int64

About the Features#

Our table includes details like:

  • Sex (male or female)
  • Age
  • Passenger class (1st, 2nd, or 3rd)
  • Number of siblings or spouses aboard
  • Fare paid

These are the 'features' we will use to help predict survival.

# Let us see which columns we have.
list(data.columns)
['PassengerId',
 'Survived',
 'Pclass',
 'Name',
 'Sex',
 'Age',
 'SibSp',
 'Parch',
 'Ticket',
 'Fare',
 'Cabin',
 'Embarked']
# Let us handle missing values in the Age column by filling them in with the average age.
data['Age'].fillna(data['Age'].mean(), inplace=True)
# We need to convert the 'Sex' column from text to numbers, since models like numbers.
data['Sex'] = data['Sex'].map({'male': 0, 'female': 1})
# Let us pick just the features we want to use.
features = ['Pclass', 'Sex', 'Age', 'SibSp', 'Fare']
X = data[features]
y = data['Survived']
# Time to split our data into a training set and a test set.
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
---------------------------------------------------------------------------
NameError                                 Traceback (most recent call last)
Cell In[8], line 2
      1 # Time to split our data into a training set and a test set.
----> 2 X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

NameError: name 'train_test_split' is not defined

Training an XGBoost Model#

Now let us build a model that predicts survival!

We will use XGBoost. XGBoost stands for Extreme Gradient Boosting, and it is very popular for competition and real-world problems.

# Here is how we fit an XGBoost model to our data.
model = xgb.XGBClassifier(use_label_encoder=False, eval_metric='logloss', random_state=42)
model.fit(X_train, y_train)
---------------------------------------------------------------------------
NameError                                 Traceback (most recent call last)
Cell In[9], line 2
      1 # Here is how we fit an XGBoost model to our data.
----> 2 model = xgb.XGBClassifier(use_label_encoder=False, eval_metric='logloss', random_state=42)
      3 model.fit(X_train, y_train)

NameError: name 'xgb' is not defined
# Let us use the trained model to predict on the test set.
y_pred = model.predict(X_test)
---------------------------------------------------------------------------
NameError                                 Traceback (most recent call last)
Cell In[10], line 2
      1 # Let us use the trained model to predict on the test set.
----> 2 y_pred = model.predict(X_test)

NameError: name 'model' is not defined
# How well did our model do?
acc = accuracy_score(y_test, y_pred)
print(f"Accuracy: {acc:.2f}")
---------------------------------------------------------------------------
NameError                                 Traceback (most recent call last)
Cell In[11], line 2
      1 # How well did our model do?
----> 2 acc = accuracy_score(y_test, y_pred)
      3 print(f"Accuracy: {acc:.2f}")

NameError: name 'accuracy_score' is not defined

XGBoost Model Parameters#

Every XGBoost model has settings we can adjust, called 'hyperparameters'.

The most important ones are:

  • n_estimators: how many trees
  • max_depth: how deep each tree goes
  • learning_rate: how much each tree corrects the last

Changing these can make your model better or worse.

# Let us try changing some model parameters.
model2 = xgb.XGBClassifier(n_estimators=50, max_depth=3, learning_rate=0.1, use_label_encoder=False, eval_metric='logloss', random_state=42)
model2.fit(X_train, y_train)
y_pred2 = model2.predict(X_test)
print("Accuracy with new params:", accuracy_score(y_test, y_pred2))
---------------------------------------------------------------------------
NameError                                 Traceback (most recent call last)
Cell In[12], line 2
      1 # Let us try changing some model parameters.
----> 2 model2 = xgb.XGBClassifier(n_estimators=50, max_depth=3, learning_rate=0.1, use_label_encoder=False, eval_metric='logloss', random_state=42)
      3 model2.fit(X_train, y_train)
      4 y_pred2 = model2.predict(X_test)

NameError: name 'xgb' is not defined
# You can ask the user to adjust parameters with input().
n_trees = int(input("Enter how many trees to use (try 10 to 200): "))
depth = int(input("Enter tree depth (try 2 to 8): "))
model3 = xgb.XGBClassifier(n_estimators=n_trees, max_depth=depth, use_label_encoder=False, eval_metric='logloss', random_state=42)
model3.fit(X_train, y_train)
y_pred3 = model3.predict(X_test)
print("Custom Accuracy:", accuracy_score(y_test, y_pred3))
---------------------------------------------------------------------------
NameError                                 Traceback (most recent call last)
Cell In[13], line 4
      2 n_trees = int(input("Enter how many trees to use (try 10 to 200): "))
      3 depth = int(input("Enter tree depth (try 2 to 8): "))
----> 4 model3 = xgb.XGBClassifier(n_estimators=n_trees, max_depth=depth, use_label_encoder=False, eval_metric='logloss', random_state=42)
      5 model3.fit(X_train, y_train)
      6 y_pred3 = model3.predict(X_test)

NameError: name 'xgb' is not defined
 

Feature Importance#

One great strength of XGBoost is showing which features matter most.

Let us look at which feature helped most in predicting survival on the Titanic!

# Get the feature importances from our model.
import matplotlib.pyplot as plt
importances = model.feature_importances_
plt.bar(features, importances)
plt.title("Feature Importances")
plt.show()
---------------------------------------------------------------------------
NameError                                 Traceback (most recent call last)
Cell In[14], line 3
      1 # Get the feature importances from our model.
      2 import matplotlib.pyplot as plt
----> 3 importances = model.feature_importances_
      4 plt.bar(features, importances)
      5 plt.title("Feature Importances")

NameError: name 'model' is not defined

Mini Project: Predict a Single Passenger#

Let us use our XGBoost model to predict for just one new Titanic passenger.

Input some details to see if they might have survived!

# We will ask for simple details then predict survival.
pclass = int(input('Enter class (1 = 1st, 2 = 2nd, 3 = 3rd): '))
sex = int(input('Enter sex (0 = male, 1 = female): '))
age = float(input('Enter age (in years): '))
sibsp = int(input('How many siblings/spouses aboard?: '))
fare = float(input('Fare paid (e.g. 25.0): '))
single_passenger = np.array([[pclass, sex, age, sibsp, fare]])
pred = model.predict(single_passenger)[0]
print('Survived!' if pred == 1 else 'Did not survive.')
---------------------------------------------------------------------------
NameError                                 Traceback (most recent call last)
Cell In[15], line 8
      6 fare = float(input('Fare paid (e.g. 25.0): '))
      7 single_passenger = np.array([[pclass, sex, age, sibsp, fare]])
----> 8 pred = model.predict(single_passenger)[0]
      9 print('Survived!' if pred == 1 else 'Did not survive.')

NameError: name 'model' is not defined
 
# Sometimes, your model will not work as expected. Let us check for missing values that may cause problems.
missing = X_test.isnull().sum()
print("Missing values:\n", missing)
---------------------------------------------------------------------------
NameError                                 Traceback (most recent call last)
Cell In[16], line 2
      1 # Sometimes, your model will not work as expected. Let us check for missing values that may cause problems.
----> 2 missing = X_test.isnull().sum()
      3 print("Missing values:\n", missing)

NameError: name 'X_test' is not defined
# Pro tip: You can combine XGBoost with other tools, like pipelines from scikit-learn.
from sklearn.pipeline import Pipeline
pipe = Pipeline([
    ('clf', xgb.XGBClassifier(use_label_encoder=False, eval_metric='logloss', random_state=42))
])
pipe.fit(X_train, y_train)
print("Pipeline accuracy:", accuracy_score(y_test, pipe.predict(X_test)))
---------------------------------------------------------------------------
NameError                                 Traceback (most recent call last)
Cell In[17], line 4
      1 # Pro tip: You can combine XGBoost with other tools, like pipelines from scikit-learn.
      2 from sklearn.pipeline import Pipeline
      3 pipe = Pipeline([
----> 4     ('clf', xgb.XGBClassifier(use_label_encoder=False, eval_metric='logloss', random_state=42))
      5 ])
      6 pipe.fit(X_train, y_train)
      7 print("Pipeline accuracy:", accuracy_score(y_test, pipe.predict(X_test)))

NameError: name 'xgb' is not defined
# Challenge: Try using a different set of features.
new_features = ['Pclass', 'Sex', 'Age']
X_new = data[new_features]
X_train_new, X_test_new, y_train_new, y_test_new = train_test_split(X_new, y, test_size=0.2, random_state=42)
model_new = xgb.XGBClassifier(use_label_encoder=False, eval_metric='logloss', random_state=42)
model_new.fit(X_train_new, y_train_new)
print("Accuracy with fewer features:", accuracy_score(y_test_new, model_new.predict(X_test_new)))
---------------------------------------------------------------------------
NameError                                 Traceback (most recent call last)
Cell In[18], line 4
      2 new_features = ['Pclass', 'Sex', 'Age']
      3 X_new = data[new_features]
----> 4 X_train_new, X_test_new, y_train_new, y_test_new = train_test_split(X_new, y, test_size=0.2, random_state=42)
      5 model_new = xgb.XGBClassifier(use_label_encoder=False, eval_metric='logloss', random_state=42)
      6 model_new.fit(X_train_new, y_train_new)

NameError: name 'train_test_split' is not defined

Recap: What Have We Learned?#

  • How to load and explore a real dataset
  • What gradient boosting is, in simple terms
  • How to train and improve an XGBoost model
  • How to check which features are important
  • Ways to test and predict survival for new passengers

You now have the basics to use XGBoost in your own projects!

Next Steps#

Try the challenge cells again, or use this template for your own dataset.

Experimenting is the best way to learn.

Subscribe to the channel for more hands-on Python notebooks and challenges!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.