Mathew K Analytics

Lesson 17 · Scikit-learn deep dive

Scikit-learn Tutorial #17: Model Persistence & Diagnostics

Video seventeen of the eighteen-part series: saving trained models and diagnosing how well they're fitting. joblib dump/load, learningcurve,…

⬇ Download notebookOpen in Colab ↗

📓 Full notebook

Download .ipynb

Scikit-learn Deep-Dive, Video 17: Model Persistence and Diagnostics#

  • Video seventeen of the eighteen-part series: saving trained models and diagnosing how well they're fitting.
  • joblib dump/load, learning_curve, validation_curve, and class_weight.
  • Let's get into it.

Part 1: joblib.dump and load Basics#

import joblib
from sklearn.datasets import load_wine
from sklearn.ensemble import RandomForestClassifier
X, y = load_wine(return_X_y=True)
model = RandomForestClassifier(random_state=42, n_estimators=100).fit(X, y)
joblib.dump(model, 'wine_model.joblib')
loaded_model = joblib.load('wine_model.joblib')
print(round(loaded_model.score(X, y), 3))
1.0

Part 2: Saving a Full Pipeline, Not Just the Model#

from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
pipe = Pipeline([
    ('scaler', StandardScaler()),
    ('model', LogisticRegression(max_iter=5000))
]).fit(X, y)
joblib.dump(pipe, 'wine_pipeline.joblib')
loaded_pipe = joblib.load('wine_pipeline.joblib')
print(round(loaded_pipe.score(X, y), 3))
1.0

Part 3: Saving with Metadata for Reproducibility#

import sklearn
bundle = {
    'model': pipe,
    'sklearn_version': sklearn.__version__,
    'feature_count': X.shape[1],
    'training_samples': X.shape[0]
}
joblib.dump(bundle, 'wine_bundle.joblib')
loaded_bundle = joblib.load('wine_bundle.joblib')
print(loaded_bundle['sklearn_version'], loaded_bundle['feature_count'])
1.7.1 13

Part 4: learning_curve - Diagnosing Under/Overfitting from Data Size#

from sklearn.model_selection import learning_curve
train_sizes, train_scores, val_scores = learning_curve(
    RandomForestClassifier(random_state=42, n_estimators=50), X, y, cv=5,
    train_sizes=[0.2, 0.4, 0.6, 0.8, 1.0]
)
print(train_sizes)
print(train_scores.mean(axis=1).round(3))
print(val_scores.mean(axis=1).round(3))
[ 28  56  85 113 142]
[1. 1. 1. 1. 1.]
[0.331 0.635 0.725 0.95  0.961]

Part 5: Interpreting the Learning Curve Gap#

gap = train_scores.mean(axis=1) - val_scores.mean(axis=1)
print(gap.round(3))
if gap[-1] > 0.05:
    print('some real overfitting signal remains, even at full data size')
else:
    print('train and validation scores are genuinely close, a healthy fit')
[0.669 0.365 0.275 0.05  0.039]
train and validation scores are genuinely close, a healthy fit

Part 6: validation_curve - Diagnosing Over/Underfitting from a Hyperparameter#

from sklearn.model_selection import validation_curve
depth_range = [1, 2, 3, 5, 8, 12, None]
train_scores_vc, val_scores_vc = validation_curve(
    RandomForestClassifier(random_state=42, n_estimators=50), X, y,
    param_name='max_depth', param_range=depth_range, cv=5
)
print(train_scores_vc.mean(axis=1).round(3))
print(val_scores_vc.mean(axis=1).round(3))
[0.985 0.987 0.996 1.    1.    1.    1.   ]
[0.961 0.972 0.972 0.967 0.961 0.961 0.961]

Part 7: class_weight='balanced' - Handling Imbalanced Data#

import numpy as np
y_imbalanced = (y == 0).astype(int)
print(np.bincount(y_imbalanced))
balanced_model = LogisticRegression(max_iter=5000, class_weight='balanced').fit(X, y_imbalanced)
print(round(balanced_model.score(X, y_imbalanced), 3))
[119  59]
0.989

Part 8: Comparing With and Without class_weight#

from sklearn.metrics import recall_score
unbalanced_model = LogisticRegression(max_iter=5000).fit(X, y_imbalanced)
unbalanced_recall = recall_score(y_imbalanced, unbalanced_model.predict(X))
balanced_recall = recall_score(y_imbalanced, balanced_model.predict(X))
print('unbalanced recall:', round(unbalanced_recall, 3))
print('balanced recall:', round(balanced_recall, 3))
unbalanced recall: 0.966
balanced recall: 1.0

Part 9: Combining class_weight with cross_val_score#

from sklearn.model_selection import cross_val_score
unbalanced_cv = cross_val_score(
    LogisticRegression(max_iter=5000), X, y_imbalanced, cv=5, scoring='recall'
)
balanced_cv = cross_val_score(
    LogisticRegression(max_iter=5000, class_weight='balanced'), X, y_imbalanced, cv=5, scoring='recall'
)
print('unbalanced CV recall:', round(unbalanced_cv.mean(), 3))
print('balanced CV recall:', round(balanced_cv.mean(), 3))
unbalanced CV recall: 0.967
balanced CV recall: 0.967

Part 10: A Real Pattern - save_model and load_model Functions#

def save_model(model, path, **metadata):
    bundle = {'model': model, 'sklearn_version': sklearn.__version__, **metadata}
    joblib.dump(bundle, path)
def load_model(path):
    bundle = joblib.load(path)
    return bundle['model'], {k: v for k, v in bundle.items() if k != 'model'}
save_model(balanced_model, 'final_model.joblib', dataset='wine_imbalanced', recall=round(balanced_recall, 3))
loaded, meta = load_model('final_model.joblib')
print(meta)
{'sklearn_version': '1.7.1', 'dataset': 'wine_imbalanced', 'recall': 1.0}

Wrap-Up: What You Learned#

  • joblib.dump and joblib.load save and restore a fitted estimator efficiently, better than plain pickle for numpy-heavy models.
  • joblib works identically on a whole fitted Pipeline, preserving every step's learned state.
  • Bundling metadata like the scikit-learn version and data shape alongside a saved model aids reproducibility.
  • learning_curve tracks train and validation scores across increasing training set sizes.
  • A persistent train/validation gap signals overfitting; both scores low and close signals underfitting.
  • validation_curve varies one hyperparameter, showing the transition from underfitting to overfitting directly.
  • class_weight='balanced' re-weights training examples inversely to class frequency, without manual resampling.
  • Comparing recall with and without class_weight shows its effect directly on the rare class.
  • Combining class_weight with cross_val_score confirms an improvement holds up under honest cross-validation.
  • That wraps up model persistence and diagnostics. Next up: the Capstone - an end-to-end production ML pipeline.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.