Lesson 29 · Python For Machine Learning
Master Ensemble Learning & Stacking Techniques in Python for Better Machine Learning Models
Welcome! Today we will explore how ensemble techniques can make your machine learning models smarter, even if you are just starting with Python. We will…
- CoursePython For Machine Learning
- Lesson29 of 16
- Video15 min
- FormatJupyter notebook · 18 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynb
Beginner Python: Ensemble Learning and Stacking#
Welcome! Today we will explore how ensemble techniques can make your machine learning models smarter, even if you are just starting with Python.
We will learn the basics step by step and try out everything in code.
By the end, you will build your own stacking ensemble using real data.
# Before we begin, let us ensure we do not see distracting warnings.
import warnings
warnings.filterwarnings("ignore")
What is Ensemble Learning?#
Instead of relying on just one model, an ensemble combines several models to boost accuracy and reduce mistakes.
A simple example is voting: each model gives its answer, and the group decides together.
# First, let us import the tools we will need for ensembles.
import numpy as np
import pandas as pd
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score
Data setup#
For this lesson, let us work with the Titanic dataset.
We will predict who survived based on their information.
# Let us load Titanic data from the internet.
url = "https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv"
df = pd.read_csv(url)
print("Shape:", df.shape)
df.head()
# Let us check how many people survived or not.
df['Survived'].value_counts()
# Let us pick a few features to make things simple.
features = ['Pclass', 'Sex', 'Age', 'Fare', 'SibSp', 'Parch']
data = df[features + ['Survived']].copy()
# Convert Sex to numbers: male = 0, female = 1
data['Sex'] = data['Sex'].map({'male': 0, 'female': 1})
# Fill missing Ages with median age
data['Age'] = data['Age'].fillna(data['Age'].median())
# Time to split into training and test data!
X = data[features]
y = data['Survived']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
Trying Basic Models#
Let us build a few starter models. These will become the parts of our ensemble later.
# A decision tree is a simple model that makes choices by asking questions.
from sklearn.tree import DecisionTreeClassifier
dt = DecisionTreeClassifier(random_state=42)
dt.fit(X_train, y_train)
dt_preds = dt.predict(X_test)
print("Decision Tree accuracy:", accuracy_score(y_test, dt_preds))
# Logistic regression is a simple way to predict yes or no by learning patterns.
from sklearn.linear_model import LogisticRegression
lr = LogisticRegression(max_iter=500)
lr.fit(X_train, y_train)
lr_preds = lr.predict(X_test)
print("Logistic Regression accuracy:", accuracy_score(y_test, lr_preds))
# Random Forest uses many trees and averages their answers.
from sklearn.ensemble import RandomForestClassifier
rf = RandomForestClassifier(n_estimators=100, random_state=42)
rf.fit(X_train, y_train)
rf_preds = rf.predict(X_test)
print("Random Forest accuracy:", accuracy_score(y_test, rf_preds))
Let us try Voting Ensemble#
Voting combines the predictions of several models. It often does better than just using one model.
# VotingClassifier lets us blend several models at once.
from sklearn.ensemble import VotingClassifier
voting = VotingClassifier(estimators=[
("logreg", lr),
("dtree", dt),
("rforest", rf)
])
voting.fit(X_train, y_train)
voting_preds = voting.predict(X_test)
print("Voting Ensemble accuracy:", accuracy_score(y_test, voting_preds))
What is Stacking?#
Stacking is a powerful kind of ensemble.
Instead of simply voting, we train a new model to learn from the predictions of other models.
This meta-model tries to learn when to trust each base model.
# Let us do stacking with StackingClassifier.
from sklearn.ensemble import StackingClassifier
stack = StackingClassifier(estimators=[
("lr", lr),
("dt", dt),
("rf", rf)
], final_estimator=LogisticRegression(max_iter=500), passthrough=False, cv=5)
stack.fit(X_train, y_train)
stack_preds = stack.predict(X_test)
print("Stacking accuracy:", accuracy_score(y_test, stack_preds))
# Mini-Project Part 1: Try stacking with a new meta-model.
from sklearn.svm import SVC
svc = SVC(probability=True, random_state=42)
stack2 = StackingClassifier(estimators=[
("lr", lr),
("dt", dt),
("rf", rf)
], final_estimator=svc, passthrough=False, cv=5)
stack2.fit(X_train, y_train)
stack2_preds = stack2.predict(X_test)
print("Stacking (SVC as meta-model) accuracy:", accuracy_score(y_test, stack2_preds))
# Let us see which models do best -- Mini-Project Part 2!
results = pd.DataFrame({
"Decision Tree": [accuracy_score(y_test, dt_preds)],
"Logistic Regression": [accuracy_score(y_test, lr_preds)],
"Random Forest": [accuracy_score(y_test, rf_preds)],
"Voting Ensemble": [accuracy_score(y_test, voting_preds)],
"Stacking (LR)": [accuracy_score(y_test, stack_preds)],
"Stacking (SVC)": [accuracy_score(y_test, stack2_preds)]
})
results.T.rename(columns={0: 'Accuracy'}).sort_values('Accuracy', ascending=False)
# Sometimes ensembles need careful balancing.
print("Class balance in test set:")
print(y_test.value_counts())
# Quick best practice: always shuffle the data before splitting.
shuffled = data.sample(frac=1, random_state=99).reset_index(drop=True)
shuffled.head()
# Challenge: Can you add an input model name and report its accuracy?
print("Which model? (dt, lr, rf, voting, stack, stack2)")
choice = input("Model code: ").strip()
model_scores = {
"dt": accuracy_score(y_test, dt_preds),
"lr": accuracy_score(y_test, lr_preds),
"rf": accuracy_score(y_test, rf_preds),
"voting": accuracy_score(y_test, voting_preds),
"stack": accuracy_score(y_test, stack_preds),
"stack2": accuracy_score(y_test, stack2_preds)
}
if choice in model_scores:
print("Accuracy:", model_scores[choice])
else:
print("Sorry, code not found.")
Recap: What youve learned#
- You discovered what an ensemble is and why it matters.
- You tried voting and stacking with simple, real data.
- You saw practical tips and best practices to start building ensembles yourself!
Challenge for you!#
Change the features, try different base models, or pick another dataset using the same steps.
Experiment and see how high your accuracy can go.
Let us keep learning together. If you enjoyed this video, give it a thumbs up, subscribe, and share your project below!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



