Lesson 6 · Scikit-learn deep dive
Scikit-learn Tutorial #6: Pipelines
Video six of the eighteen-part series: chaining preprocessing and a model into one real estimator. Pipeline, makepipeline, named steps, and avoiding data…
- CourseScikit-learn deep dive
- Lesson6 of 18
- Video15 min
- FormatJupyter notebook · 10 code cells
What you'll learn
- Why Pipelines - the Leakage Problem
- Building a Pipeline with Named Steps
- makepipeline - Auto-Naming
- fit/predict/score - Treating the Whole Thing as One Estimator
- Accessing Steps - namedsteps, Indexing, Slicing
- Pipeline Parameter Naming - stepparam
- Chaining Multiple Preprocessing Steps
- Avoiding Leakage with crossvalscore
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbScikit-learn Deep-Dive, Video 6: Pipelines#
- Video six of the eighteen-part series: chaining preprocessing and a model into one real estimator.
- Pipeline, make_pipeline, named steps, and avoiding data leakage.
- Let's get into it.
Part 1: Why Pipelines - the Leakage Problem#
import numpy as np
from sklearn.datasets import load_wine
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
X, y = load_wine(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, random_state=0)
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)
X_test_scaled = scaler.transform(X_test)
Part 2: Building a Pipeline with Named Steps#
from sklearn.pipeline import Pipeline
pipe = Pipeline([
('scaler', StandardScaler()),
('model', LogisticRegression(max_iter=1000))
])
pipe.fit(X_train, y_train)
print(pipe.named_steps.keys())
Part 3: make_pipeline - Auto-Naming#
from sklearn.pipeline import make_pipeline
auto_pipe = make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000))
print(list(auto_pipe.named_steps.keys()))
auto_pipe.fit(X_train, y_train)
print(round(auto_pipe.score(X_test, y_test), 3))
Part 4: fit/predict/score - Treating the Whole Thing as One Estimator#
predictions = pipe.predict(X_test)
print(predictions[:10])
print(y_test[:10])
print(round(pipe.score(X_test, y_test), 3))
Part 5: Accessing Steps - named_steps, Indexing, Slicing#
fitted_scaler = pipe.named_steps['scaler']
print(fitted_scaler.mean_.round(2))
print(pipe[0])
print(pipe[-1].coef_.shape)
print(pipe[:1])
Part 6: Pipeline Parameter Naming - step__param#
print(sorted(pipe.get_params().keys())[:6])
pipe.set_params(model__C=0.5)
print(pipe.named_steps['model'].C)
Part 7: Chaining Multiple Preprocessing Steps#
from sklearn.impute import SimpleImputer
X_with_gaps = X_train.copy()
X_with_gaps[0, 0] = np.nan
full_pipe = Pipeline([
('imputer', SimpleImputer(strategy='mean')),
('scaler', StandardScaler()),
('model', LogisticRegression(max_iter=1000))
])
full_pipe.fit(X_with_gaps, y_train)
print(round(full_pipe.score(X_test, y_test), 3))
Part 8: Avoiding Leakage with cross_val_score#
from sklearn.model_selection import cross_val_score
scores = cross_val_score(pipe, X, y, cv=5)
print(scores.round(3))
print(round(scores.mean(), 3))
Part 9: The Wrong Way vs the Right Way, Concretely#
leaky_scaler = StandardScaler().fit(X)
X_leaky = leaky_scaler.transform(X)
from sklearn.linear_model import LogisticRegression as LR2
leaky_scores = cross_val_score(LR2(max_iter=1000), X_leaky, y, cv=5)
safe_scores = cross_val_score(pipe, X, y, cv=5)
print(round(leaky_scores.mean(), 3), round(safe_scores.mean(), 3))
Part 10: A Real Pattern - a Reusable build_pipeline Function#
def build_pipeline(model, impute_strategy='mean'):
return Pipeline([
('imputer', SimpleImputer(strategy=impute_strategy)),
('scaler', StandardScaler()),
('model', model)
])
candidate_pipe = build_pipeline(LogisticRegression(max_iter=1000))
candidate_scores = cross_val_score(candidate_pipe, X, y, cv=5)
print(round(candidate_scores.mean(), 3))
Wrap-Up: What You Learned#
- Manually fitting and transforming preprocessing steps separately invites leakage from accidentally re-fitting on test data.
- Pipeline bundles named steps, transformers then a final estimator, into one object with a single fit call.
- make_pipeline builds the same structure with auto-generated step names.
- A fitted pipeline exposes the same fit/predict/score interface as any single estimator.
- named_steps, indexing, and slicing all reach individual fitted steps for inspection.
- step__param naming reaches parameters deep inside a pipeline, the syntax GridSearchCV relies on.
- Pipelines can chain multiple preprocessing steps, like an imputer then a scaler, before the final model.
- Passing a whole pipeline into cross_val_score keeps every fold's preprocessing correctly isolated.
- Scaling the full dataset before splitting is the classic leakage mistake; pipelines make it structurally impossible.
- That wraps up pipelines. Next up: ColumnTransformer, applying different preprocessing per column type.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



