Lesson 18 · Scikit-learn deep dive
Scikit-learn Tutorial #18: Capstone — End-to-End ML Pipeline
The final video of the eighteen-part series: combining everything into one real workflow. ColumnTransformer, Pipeline, GridSearchCV, metrics, and joblib…
- CourseScikit-learn deep dive
- Lesson18 of 18
- Video16 min
- FormatJupyter notebook · 10 code cells
What you'll learn
- The Full Picture - Recapping the Whole Series
- Building a Realistic Mixed Dataset with Missing Values
- train/test Split with Stratification
- Building the ColumnTransformer
- Wrapping in a Full Pipeline with a Classifier
- Tuning with GridSearchCV
- Evaluating with classificationreport and ROC-AUC
- crossvalscore Sanity Check on the Full Pipeline
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbScikit-learn Deep-Dive, Video 18: Capstone - End-to-End Production ML Pipeline#
- The final video of the eighteen-part series: combining everything into one real workflow.
- ColumnTransformer, Pipeline, GridSearchCV, metrics, and joblib persistence, together.
- Let's get into it.
Part 1: The Full Picture - Recapping the Whole Series#
import numpy as np
import pandas as pd
print('starting the real end-to-end capstone workflow')
Part 2: Building a Realistic Mixed Dataset with Missing Values#
rng = np.random.RandomState(42)
n = 400
df = pd.DataFrame({
'tenure_months': rng.randint(1, 72, n).astype(float),
'monthly_charge': rng.normal(65, 20, n).round(2),
'support_calls': rng.poisson(2, n).astype(float),
'contract_type': rng.choice(['month-to-month', 'one-year', 'two-year'], n),
'internet_service': rng.choice(['dsl', 'fiber', 'none'], n)
})
df.loc[rng.choice(n, 30, replace=False), 'monthly_charge'] = np.nan
df.loc[rng.choice(n, 20, replace=False), 'internet_service'] = None
churn_score = (
(df['contract_type'] == 'month-to-month').astype(int) * 2
+ (df['support_calls'] > 3).astype(int) * 2
+ (df['tenure_months'] < 12).astype(int)
+ rng.normal(0, 1, n)
)
y = (churn_score > churn_score.median()).astype(int)
print(df.shape, df.isna().sum().sum())
Part 3: train/test Split with Stratification#
from sklearn.model_selection import train_test_split
X = df.copy()
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42, stratify=y
)
print(X_train.shape, X_test.shape)
Part 4: Building the ColumnTransformer#
from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.impute import SimpleImputer
from sklearn.preprocessing import StandardScaler, OneHotEncoder
numeric_cols = ['tenure_months', 'monthly_charge', 'support_calls']
categorical_cols = ['contract_type', 'internet_service']
numeric_pipe = Pipeline([
('impute', SimpleImputer(strategy='median')),
('scale', StandardScaler())
])
categorical_pipe = Pipeline([
('impute', SimpleImputer(strategy='most_frequent')),
('encode', OneHotEncoder(handle_unknown='ignore'))
])
preprocessor = ColumnTransformer([
('num', numeric_pipe, numeric_cols),
('cat', categorical_pipe, categorical_cols)
])
Part 5: Wrapping in a Full Pipeline with a Classifier#
from sklearn.ensemble import RandomForestClassifier
full_pipeline = Pipeline([
('preprocess', preprocessor),
('model', RandomForestClassifier(random_state=42))
])
full_pipeline.fit(X_train, y_train)
print(round(full_pipeline.score(X_test, y_test), 3))
Part 6: Tuning with GridSearchCV#
from sklearn.model_selection import GridSearchCV
param_grid = {
'model__n_estimators': [100, 200],
'model__max_depth': [5, 10, None]
}
search = GridSearchCV(full_pipeline, param_grid, cv=5, scoring='f1')
search.fit(X_train, y_train)
print(search.best_params_)
print(round(search.best_score_, 3))
Part 7: Evaluating with classification_report and ROC-AUC#
from sklearn.metrics import classification_report, roc_auc_score
best_model = search.best_estimator_
test_preds = best_model.predict(X_test)
test_probs = best_model.predict_proba(X_test)[:, 1]
print(classification_report(y_test, test_preds, target_names=['stayed', 'churned']))
print(round(roc_auc_score(y_test, test_probs), 3))
Part 8: cross_val_score Sanity Check on the Full Pipeline#
from sklearn.model_selection import cross_val_score
sanity_scores = cross_val_score(best_model, X, y, cv=5, scoring='f1')
print(sanity_scores.round(3))
print(round(sanity_scores.mean(), 3))
Part 9: Persisting the Tuned Pipeline with joblib#
import joblib
import sklearn
bundle = {
'model': best_model,
'sklearn_version': sklearn.__version__,
'best_params': search.best_params_,
'test_roc_auc': round(roc_auc_score(y_test, test_probs), 3)
}
joblib.dump(bundle, 'churn_model_bundle.joblib')
print('saved successfully')
Part 10: Loading It Back and Predicting on Brand-New Raw Data#
loaded_bundle = joblib.load('churn_model_bundle.joblib')
production_model = loaded_bundle['model']
new_customers = pd.DataFrame({
'tenure_months': [2.0, 48.0, np.nan],
'monthly_charge': [89.99, 45.00, 60.0],
'support_calls': [5.0, 0.0, 1.0],
'contract_type': ['month-to-month', 'two-year', 'one-year'],
'internet_service': ['fiber', 'dsl', None]
})
predictions = production_model.predict(new_customers)
probabilities = production_model.predict_proba(new_customers)[:, 1]
print(list(zip(predictions, probabilities.round(3))))
Wrap-Up: What You Learned#
- This capstone combined every piece from the series into one real, production-shaped workflow.
- A realistic dataset mixes numeric and categorical columns with genuine missing values.
- Splitting before anything else, with stratification, is always the honest first step.
- ColumnTransformer combines a numeric impute-then-scale pipeline with a categorical impute-then-encode pipeline.
- The full preprocessor plus classifier becomes one Pipeline, handling raw data straight through to a prediction.
- GridSearchCV with step__param naming tunes the model while keeping preprocessing leakage-safe inside every fold.
- classification_report and ROC-AUC on the held-out test set give an honest, complete performance picture.
- A final independent cross_val_score pass sanity-checks that performance isn't just a lucky split.
- joblib.dump saves the whole tuned pipeline plus metadata, making the workflow genuinely production-ready.
- Loading it back predicts directly on brand-new raw data, no manual preprocessing step required, completing the full circle.
- That's the full eighteen-part scikit-learn deep-dive series, complete. Thanks for following along the whole way.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



