Lesson 15 · Scikit-learn deep dive
Scikit-learn Tutorial #15: Ensembling & Meta-Estimators
Video fifteen of the eighteen-part series: combining multiple models into one stronger predictor. VotingClassifier, BaggingClassifier, and…
- CourseScikit-learn deep dive
- Lesson15 of 18
- Video14 min
- FormatJupyter notebook · 10 code cells
What you'll learn
- Why Ensembles - Combining Models Often Beats Any Single One
- VotingClassifier - Hard Voting
- VotingClassifier - Soft Voting
- weights - Giving Stronger Models More Say
- BaggingClassifier - Bootstrap Aggregating
- BaggingClassifier - oobscore
- StackingClassifier - a Meta-Learner on Base Predictions
- Comparing Individual Models vs Ensembles with crossvalscore
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbScikit-learn Deep-Dive, Video 15: Ensembling Meta-Estimators#
- Video fifteen of the eighteen-part series: combining multiple models into one stronger predictor.
- VotingClassifier, BaggingClassifier, and StackingClassifier.
- Let's get into it.
Part 1: Why Ensembles - Combining Models Often Beats Any Single One#
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split
X, y = load_breast_cancer(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, random_state=42, stratify=y)
print(X_train.shape, X_test.shape)
Part 2: VotingClassifier - Hard Voting#
from sklearn.ensemble import VotingClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.tree import DecisionTreeClassifier
from sklearn.neighbors import KNeighborsClassifier
voters = [
('lr', LogisticRegression(max_iter=5000)),
('dt', DecisionTreeClassifier(random_state=42, max_depth=4)),
('knn', KNeighborsClassifier(n_neighbors=5))
]
hard_voter = VotingClassifier(estimators=voters, voting='hard')
hard_voter.fit(X_train, y_train)
print(round(hard_voter.score(X_test, y_test), 3))
Part 3: VotingClassifier - Soft Voting#
soft_voter = VotingClassifier(estimators=voters, voting='soft')
soft_voter.fit(X_train, y_train)
print(round(soft_voter.score(X_test, y_test), 3))
print(soft_voter.predict_proba(X_test[:3]).round(3))
Part 4: weights - Giving Stronger Models More Say#
weighted_voter = VotingClassifier(estimators=voters, voting='soft', weights=[2, 1, 1])
weighted_voter.fit(X_train, y_train)
print(round(weighted_voter.score(X_test, y_test), 3))
Part 5: BaggingClassifier - Bootstrap Aggregating#
from sklearn.ensemble import BaggingClassifier
bagger = BaggingClassifier(
DecisionTreeClassifier(random_state=42), n_estimators=50, random_state=42
)
bagger.fit(X_train, y_train)
print(round(bagger.score(X_test, y_test), 3))
Part 6: BaggingClassifier - oob_score#
oob_bagger = BaggingClassifier(
DecisionTreeClassifier(random_state=42), n_estimators=50, oob_score=True, random_state=42
)
oob_bagger.fit(X_train, y_train)
print(round(oob_bagger.oob_score_, 3))
print(round(oob_bagger.score(X_test, y_test), 3))
Part 7: StackingClassifier - a Meta-Learner on Base Predictions#
from sklearn.ensemble import StackingClassifier
stacker = StackingClassifier(
estimators=voters, final_estimator=LogisticRegression(max_iter=5000), cv=5
)
stacker.fit(X_train, y_train)
print(round(stacker.score(X_test, y_test), 3))
Part 8: Comparing Individual Models vs Ensembles with cross_val_score#
from sklearn.model_selection import cross_val_score
candidates = {
'logistic': LogisticRegression(max_iter=5000),
'tree': DecisionTreeClassifier(random_state=42, max_depth=4),
'knn': KNeighborsClassifier(n_neighbors=5),
'hard_vote': hard_voter,
'soft_vote': soft_voter,
'bagging': bagger
}
for name, model in candidates.items():
scores = cross_val_score(model, X, y, cv=5)
print(name, round(scores.mean(), 3))
Part 9: AdaBoostClassifier - a Different Ensembling Strategy#
from sklearn.ensemble import AdaBoostClassifier
booster = AdaBoostClassifier(n_estimators=50, random_state=42)
booster.fit(X_train, y_train)
print(round(booster.score(X_test, y_test), 3))
Part 10: A Real Pattern - a Reusable build_ensemble Function#
def build_ensemble(named_models, voting='soft'):
return VotingClassifier(estimators=named_models, voting=voting)
ensemble = build_ensemble(voters)
ensemble.fit(X_train, y_train)
print(round(ensemble.score(X_test, y_test), 3))
Wrap-Up: What You Learned#
- Different model types make different mistakes; combining several often cancels out individual weaknesses.
- VotingClassifier with voting='hard' predicts whichever class the majority of base models chose.
- voting='soft' averages predicted probabilities instead, using more information than a plain majority vote.
- weights lets a stronger base model count for more in the final vote or averaged probability.
- BaggingClassifier trains many copies of the same base model on bootstrap samples, mainly reducing variance.
- oob_score=True evaluates each tree on its own left-out data, giving a free validation estimate.
- StackingClassifier trains a meta-learner on base models' out-of-fold predictions instead of a fixed combining rule.
- Cross-validating individual models alongside ensembles on identical folds gives a fair comparison.
- AdaBoostClassifier trains estimators sequentially, focusing on prior mistakes, a boosting strategy reducing bias.
- That wraps up ensembling meta-estimators. Next up: Custom Estimators - BaseEstimator and TransformerMixin.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



