Lesson 17 · Scikit-learn deep dive
Scikit-learn Tutorial #17: Model Persistence & Diagnostics
Video seventeen of the eighteen-part series: saving trained models and diagnosing how well they're fitting. joblib dump/load, learningcurve,…
- CourseScikit-learn deep dive
- Lesson17 of 18
- Video15 min
- FormatJupyter notebook · 10 code cells
What you'll learn
- joblib.dump and load Basics
- Saving a Full Pipeline, Not Just the Model
- Saving with Metadata for Reproducibility
- learningcurve - Diagnosing Under/Overfitting from Data Size
- Interpreting the Learning Curve Gap
- validationcurve - Diagnosing Over/Underfitting from a Hyperparameter
- classweight='balanced' - Handling Imbalanced Data
- Comparing With and Without classweight
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbScikit-learn Deep-Dive, Video 17: Model Persistence and Diagnostics#
- Video seventeen of the eighteen-part series: saving trained models and diagnosing how well they're fitting.
- joblib dump/load, learning_curve, validation_curve, and class_weight.
- Let's get into it.
Part 1: joblib.dump and load Basics#
import joblib
from sklearn.datasets import load_wine
from sklearn.ensemble import RandomForestClassifier
X, y = load_wine(return_X_y=True)
model = RandomForestClassifier(random_state=42, n_estimators=100).fit(X, y)
joblib.dump(model, 'wine_model.joblib')
loaded_model = joblib.load('wine_model.joblib')
print(round(loaded_model.score(X, y), 3))
Part 2: Saving a Full Pipeline, Not Just the Model#
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
pipe = Pipeline([
('scaler', StandardScaler()),
('model', LogisticRegression(max_iter=5000))
]).fit(X, y)
joblib.dump(pipe, 'wine_pipeline.joblib')
loaded_pipe = joblib.load('wine_pipeline.joblib')
print(round(loaded_pipe.score(X, y), 3))
Part 3: Saving with Metadata for Reproducibility#
import sklearn
bundle = {
'model': pipe,
'sklearn_version': sklearn.__version__,
'feature_count': X.shape[1],
'training_samples': X.shape[0]
}
joblib.dump(bundle, 'wine_bundle.joblib')
loaded_bundle = joblib.load('wine_bundle.joblib')
print(loaded_bundle['sklearn_version'], loaded_bundle['feature_count'])
Part 4: learning_curve - Diagnosing Under/Overfitting from Data Size#
from sklearn.model_selection import learning_curve
train_sizes, train_scores, val_scores = learning_curve(
RandomForestClassifier(random_state=42, n_estimators=50), X, y, cv=5,
train_sizes=[0.2, 0.4, 0.6, 0.8, 1.0]
)
print(train_sizes)
print(train_scores.mean(axis=1).round(3))
print(val_scores.mean(axis=1).round(3))
Part 5: Interpreting the Learning Curve Gap#
gap = train_scores.mean(axis=1) - val_scores.mean(axis=1)
print(gap.round(3))
if gap[-1] > 0.05:
print('some real overfitting signal remains, even at full data size')
else:
print('train and validation scores are genuinely close, a healthy fit')
Part 6: validation_curve - Diagnosing Over/Underfitting from a Hyperparameter#
from sklearn.model_selection import validation_curve
depth_range = [1, 2, 3, 5, 8, 12, None]
train_scores_vc, val_scores_vc = validation_curve(
RandomForestClassifier(random_state=42, n_estimators=50), X, y,
param_name='max_depth', param_range=depth_range, cv=5
)
print(train_scores_vc.mean(axis=1).round(3))
print(val_scores_vc.mean(axis=1).round(3))
Part 7: class_weight='balanced' - Handling Imbalanced Data#
import numpy as np
y_imbalanced = (y == 0).astype(int)
print(np.bincount(y_imbalanced))
balanced_model = LogisticRegression(max_iter=5000, class_weight='balanced').fit(X, y_imbalanced)
print(round(balanced_model.score(X, y_imbalanced), 3))
Part 8: Comparing With and Without class_weight#
from sklearn.metrics import recall_score
unbalanced_model = LogisticRegression(max_iter=5000).fit(X, y_imbalanced)
unbalanced_recall = recall_score(y_imbalanced, unbalanced_model.predict(X))
balanced_recall = recall_score(y_imbalanced, balanced_model.predict(X))
print('unbalanced recall:', round(unbalanced_recall, 3))
print('balanced recall:', round(balanced_recall, 3))
Part 9: Combining class_weight with cross_val_score#
from sklearn.model_selection import cross_val_score
unbalanced_cv = cross_val_score(
LogisticRegression(max_iter=5000), X, y_imbalanced, cv=5, scoring='recall'
)
balanced_cv = cross_val_score(
LogisticRegression(max_iter=5000, class_weight='balanced'), X, y_imbalanced, cv=5, scoring='recall'
)
print('unbalanced CV recall:', round(unbalanced_cv.mean(), 3))
print('balanced CV recall:', round(balanced_cv.mean(), 3))
Part 10: A Real Pattern - save_model and load_model Functions#
def save_model(model, path, **metadata):
bundle = {'model': model, 'sklearn_version': sklearn.__version__, **metadata}
joblib.dump(bundle, path)
def load_model(path):
bundle = joblib.load(path)
return bundle['model'], {k: v for k, v in bundle.items() if k != 'model'}
save_model(balanced_model, 'final_model.joblib', dataset='wine_imbalanced', recall=round(balanced_recall, 3))
loaded, meta = load_model('final_model.joblib')
print(meta)
Wrap-Up: What You Learned#
- joblib.dump and joblib.load save and restore a fitted estimator efficiently, better than plain pickle for numpy-heavy models.
- joblib works identically on a whole fitted Pipeline, preserving every step's learned state.
- Bundling metadata like the scikit-learn version and data shape alongside a saved model aids reproducibility.
- learning_curve tracks train and validation scores across increasing training set sizes.
- A persistent train/validation gap signals overfitting; both scores low and close signals underfitting.
- validation_curve varies one hyperparameter, showing the transition from underfitting to overfitting directly.
- class_weight='balanced' re-weights training examples inversely to class frequency, without manual resampling.
- Comparing recall with and without class_weight shows its effect directly on the rare class.
- Combining class_weight with cross_val_score confirms an improvement holds up under honest cross-validation.
- That wraps up model persistence and diagnostics. Next up: the Capstone - an end-to-end production ML pipeline.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



