Mathew K Analytics

Lesson 18 · Scikit-learn deep dive

Scikit-learn Tutorial #18: Capstone — End-to-End ML Pipeline

The final video of the eighteen-part series: combining everything into one real workflow. ColumnTransformer, Pipeline, GridSearchCV, metrics, and joblib…

⬇ Download notebookOpen in Colab ↗

📓 Full notebook

Download .ipynb

Scikit-learn Deep-Dive, Video 18: Capstone - End-to-End Production ML Pipeline#

  • The final video of the eighteen-part series: combining everything into one real workflow.
  • ColumnTransformer, Pipeline, GridSearchCV, metrics, and joblib persistence, together.
  • Let's get into it.

Part 1: The Full Picture - Recapping the Whole Series#

import numpy as np
import pandas as pd
print('starting the real end-to-end capstone workflow')
starting the real end-to-end capstone workflow

Part 2: Building a Realistic Mixed Dataset with Missing Values#

rng = np.random.RandomState(42)
n = 400
df = pd.DataFrame({
    'tenure_months': rng.randint(1, 72, n).astype(float),
    'monthly_charge': rng.normal(65, 20, n).round(2),
    'support_calls': rng.poisson(2, n).astype(float),
    'contract_type': rng.choice(['month-to-month', 'one-year', 'two-year'], n),
    'internet_service': rng.choice(['dsl', 'fiber', 'none'], n)
})
df.loc[rng.choice(n, 30, replace=False), 'monthly_charge'] = np.nan
df.loc[rng.choice(n, 20, replace=False), 'internet_service'] = None
churn_score = (
    (df['contract_type'] == 'month-to-month').astype(int) * 2
    + (df['support_calls'] > 3).astype(int) * 2
    + (df['tenure_months'] < 12).astype(int)
    + rng.normal(0, 1, n)
)
y = (churn_score > churn_score.median()).astype(int)
print(df.shape, df.isna().sum().sum())
(400, 5) 50

Part 3: train/test Split with Stratification#

from sklearn.model_selection import train_test_split
X = df.copy()
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42, stratify=y
)
print(X_train.shape, X_test.shape)
(320, 5) (80, 5)

Part 4: Building the ColumnTransformer#

from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.impute import SimpleImputer
from sklearn.preprocessing import StandardScaler, OneHotEncoder
numeric_cols = ['tenure_months', 'monthly_charge', 'support_calls']
categorical_cols = ['contract_type', 'internet_service']
numeric_pipe = Pipeline([
    ('impute', SimpleImputer(strategy='median')),
    ('scale', StandardScaler())
])
categorical_pipe = Pipeline([
    ('impute', SimpleImputer(strategy='most_frequent')),
    ('encode', OneHotEncoder(handle_unknown='ignore'))
])
preprocessor = ColumnTransformer([
    ('num', numeric_pipe, numeric_cols),
    ('cat', categorical_pipe, categorical_cols)
])

Part 5: Wrapping in a Full Pipeline with a Classifier#

from sklearn.ensemble import RandomForestClassifier
full_pipeline = Pipeline([
    ('preprocess', preprocessor),
    ('model', RandomForestClassifier(random_state=42))
])
full_pipeline.fit(X_train, y_train)
print(round(full_pipeline.score(X_test, y_test), 3))
0.825

Part 6: Tuning with GridSearchCV#

from sklearn.model_selection import GridSearchCV
param_grid = {
    'model__n_estimators': [100, 200],
    'model__max_depth': [5, 10, None]
}
search = GridSearchCV(full_pipeline, param_grid, cv=5, scoring='f1')
search.fit(X_train, y_train)
print(search.best_params_)
print(round(search.best_score_, 3))
{'model__max_depth': 10, 'model__n_estimators': 200}
0.785

Part 7: Evaluating with classification_report and ROC-AUC#

from sklearn.metrics import classification_report, roc_auc_score
best_model = search.best_estimator_
test_preds = best_model.predict(X_test)
test_probs = best_model.predict_proba(X_test)[:, 1]
print(classification_report(y_test, test_preds, target_names=['stayed', 'churned']))
print(round(roc_auc_score(y_test, test_probs), 3))
              precision    recall  f1-score   support

      stayed       0.81      0.88      0.84        40
     churned       0.86      0.80      0.83        40

    accuracy                           0.84        80
   macro avg       0.84      0.84      0.84        80
weighted avg       0.84      0.84      0.84        80

0.876

Part 8: cross_val_score Sanity Check on the Full Pipeline#

from sklearn.model_selection import cross_val_score
sanity_scores = cross_val_score(best_model, X, y, cv=5, scoring='f1')
print(sanity_scores.round(3))
print(round(sanity_scores.mean(), 3))
[0.816 0.75  0.822 0.835 0.779]
0.8

Part 9: Persisting the Tuned Pipeline with joblib#

import joblib
import sklearn
bundle = {
    'model': best_model,
    'sklearn_version': sklearn.__version__,
    'best_params': search.best_params_,
    'test_roc_auc': round(roc_auc_score(y_test, test_probs), 3)
}
joblib.dump(bundle, 'churn_model_bundle.joblib')
print('saved successfully')
saved successfully

Part 10: Loading It Back and Predicting on Brand-New Raw Data#

loaded_bundle = joblib.load('churn_model_bundle.joblib')
production_model = loaded_bundle['model']
new_customers = pd.DataFrame({
    'tenure_months': [2.0, 48.0, np.nan],
    'monthly_charge': [89.99, 45.00, 60.0],
    'support_calls': [5.0, 0.0, 1.0],
    'contract_type': ['month-to-month', 'two-year', 'one-year'],
    'internet_service': ['fiber', 'dsl', None]
})
predictions = production_model.predict(new_customers)
probabilities = production_model.predict_proba(new_customers)[:, 1]
print(list(zip(predictions, probabilities.round(3))))
[(np.int64(1), np.float64(0.94)), (np.int64(0), np.float64(0.3)), (np.int64(0), np.float64(0.053))]

Wrap-Up: What You Learned#

  • This capstone combined every piece from the series into one real, production-shaped workflow.
  • A realistic dataset mixes numeric and categorical columns with genuine missing values.
  • Splitting before anything else, with stratification, is always the honest first step.
  • ColumnTransformer combines a numeric impute-then-scale pipeline with a categorical impute-then-encode pipeline.
  • The full preprocessor plus classifier becomes one Pipeline, handling raw data straight through to a prediction.
  • GridSearchCV with step__param naming tunes the model while keeping preprocessing leakage-safe inside every fold.
  • classification_report and ROC-AUC on the held-out test set give an honest, complete performance picture.
  • A final independent cross_val_score pass sanity-checks that performance isn't just a lucky split.
  • joblib.dump saves the whole tuned pipeline plus metadata, making the workflow genuinely production-ready.
  • Loading it back predicts directly on brand-new raw data, no manual preprocessing step required, completing the full circle.
  • That's the full eighteen-part scikit-learn deep-dive series, complete. Thanks for following along the whole way.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.