Mathew K Analytics

Lesson 6 · Scikit-learn deep dive

Scikit-learn Tutorial #6: Pipelines

Video six of the eighteen-part series: chaining preprocessing and a model into one real estimator. Pipeline, makepipeline, named steps, and avoiding data…

⬇ Download notebookOpen in Colab ↗

📓 Full notebook

Download .ipynb

Scikit-learn Deep-Dive, Video 6: Pipelines#

  • Video six of the eighteen-part series: chaining preprocessing and a model into one real estimator.
  • Pipeline, make_pipeline, named steps, and avoiding data leakage.
  • Let's get into it.

Part 1: Why Pipelines - the Leakage Problem#

import numpy as np
from sklearn.datasets import load_wine
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
X, y = load_wine(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, random_state=0)
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)
X_test_scaled = scaler.transform(X_test)

Part 2: Building a Pipeline with Named Steps#

from sklearn.pipeline import Pipeline
pipe = Pipeline([
    ('scaler', StandardScaler()),
    ('model', LogisticRegression(max_iter=1000))
])
pipe.fit(X_train, y_train)
print(pipe.named_steps.keys())
dict_keys(['scaler', 'model'])

Part 3: make_pipeline - Auto-Naming#

from sklearn.pipeline import make_pipeline
auto_pipe = make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000))
print(list(auto_pipe.named_steps.keys()))
auto_pipe.fit(X_train, y_train)
print(round(auto_pipe.score(X_test, y_test), 3))
['standardscaler', 'logisticregression']
1.0

Part 4: fit/predict/score - Treating the Whole Thing as One Estimator#

predictions = pipe.predict(X_test)
print(predictions[:10])
print(y_test[:10])
print(round(pipe.score(X_test, y_test), 3))
[0 2 1 0 1 1 0 2 1 1]
[0 2 1 0 1 1 0 2 1 1]
1.0

Part 5: Accessing Steps - named_steps, Indexing, Slicing#

fitted_scaler = pipe.named_steps['scaler']
print(fitted_scaler.mean_.round(2))
print(pipe[0])
print(pipe[-1].coef_.shape)
print(pipe[:1])
[1.3000e+01 2.3900e+00 2.3700e+00 1.9510e+01 1.0046e+02 2.2600e+00
 1.9600e+00 3.6000e-01 1.6100e+00 5.1100e+00 9.5000e-01 2.5900e+00
 7.4981e+02]
StandardScaler()
(3, 13)
Pipeline(steps=[('scaler', StandardScaler())])

Part 6: Pipeline Parameter Naming - step__param#

print(sorted(pipe.get_params().keys())[:6])
pipe.set_params(model__C=0.5)
print(pipe.named_steps['model'].C)
['memory', 'model', 'model__C', 'model__class_weight', 'model__dual', 'model__fit_intercept']
0.5

Part 7: Chaining Multiple Preprocessing Steps#

from sklearn.impute import SimpleImputer
X_with_gaps = X_train.copy()
X_with_gaps[0, 0] = np.nan
full_pipe = Pipeline([
    ('imputer', SimpleImputer(strategy='mean')),
    ('scaler', StandardScaler()),
    ('model', LogisticRegression(max_iter=1000))
])
full_pipe.fit(X_with_gaps, y_train)
print(round(full_pipe.score(X_test, y_test), 3))
1.0

Part 8: Avoiding Leakage with cross_val_score#

from sklearn.model_selection import cross_val_score
scores = cross_val_score(pipe, X, y, cv=5)
print(scores.round(3))
print(round(scores.mean(), 3))
[0.972 0.972 1.    1.    1.   ]
0.989

Part 9: The Wrong Way vs the Right Way, Concretely#

leaky_scaler = StandardScaler().fit(X)
X_leaky = leaky_scaler.transform(X)
from sklearn.linear_model import LogisticRegression as LR2
leaky_scores = cross_val_score(LR2(max_iter=1000), X_leaky, y, cv=5)
safe_scores = cross_val_score(pipe, X, y, cv=5)
print(round(leaky_scores.mean(), 3), round(safe_scores.mean(), 3))
0.989 0.989

Part 10: A Real Pattern - a Reusable build_pipeline Function#

def build_pipeline(model, impute_strategy='mean'):
    return Pipeline([
        ('imputer', SimpleImputer(strategy=impute_strategy)),
        ('scaler', StandardScaler()),
        ('model', model)
    ])
candidate_pipe = build_pipeline(LogisticRegression(max_iter=1000))
candidate_scores = cross_val_score(candidate_pipe, X, y, cv=5)
print(round(candidate_scores.mean(), 3))
0.983

Wrap-Up: What You Learned#

  • Manually fitting and transforming preprocessing steps separately invites leakage from accidentally re-fitting on test data.
  • Pipeline bundles named steps, transformers then a final estimator, into one object with a single fit call.
  • make_pipeline builds the same structure with auto-generated step names.
  • A fitted pipeline exposes the same fit/predict/score interface as any single estimator.
  • named_steps, indexing, and slicing all reach individual fitted steps for inspection.
  • step__param naming reaches parameters deep inside a pipeline, the syntax GridSearchCV relies on.
  • Pipelines can chain multiple preprocessing steps, like an imputer then a scaler, before the final model.
  • Passing a whole pipeline into cross_val_score keeps every fold's preprocessing correctly isolated.
  • Scaling the full dataset before splitting is the classic leakage mistake; pipelines make it structurally impossible.
  • That wraps up pipelines. Next up: ColumnTransformer, applying different preprocessing per column type.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.