Lesson 1 · Scikit-learn deep dive
Scikit-learn Tutorial #1: The Estimator API Explained
Video one of an eighteen-part series on scikit-learn's own toolkit, not the algorithms, the actual machinery that makes it a real production library. fit,…
- CourseScikit-learn deep dive
- Lesson1 of 18
- Video17 min
- FormatJupyter notebook · 10 code cells
What you'll learn
- Why a Consistent API Matters
- Estimators and fit() - the Universal Pattern
- Predictors - predict() and predictproba()
- Transformers - transform() and fittransform()
- Fitted Attributes - the Trailing Underscore Convention
- getparams() and setparams()
- Cloning Estimators with sklearn.base.clone
- score() - the Universal Evaluation Method
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbScikit-learn Deep-Dive, Video 1: The Estimator API#
- Video one of an eighteen-part series on scikit-learn's own toolkit, not the algorithms, the actual machinery that makes it a real production library.
- fit, predict, transform, fitted attributes, get_params/set_params, cloning, and score.
- Let's get into it.
Part 1: Why a Consistent API Matters#
from sklearn.linear_model import LogisticRegression
from sklearn.tree import DecisionTreeClassifier
from sklearn.datasets import load_iris
X, y = load_iris(return_X_y=True)
model_a = LogisticRegression(max_iter=200)
model_b = DecisionTreeClassifier()
model_a.fit(X, y)
model_b.fit(X, y)
print(model_a.predict(X[:3]))
print(model_b.predict(X[:3]))
Part 2: Estimators and fit() - the Universal Pattern#
model = LogisticRegression(max_iter=200)
print(hasattr(model, 'coef_'))
returned = model.fit(X, y)
print(returned is model)
print(hasattr(model, 'coef_'))
Part 3: Predictors - predict() and predict_proba()#
model = LogisticRegression(max_iter=200).fit(X, y)
preds = model.predict(X[:5])
print(preds)
probs = model.predict_proba(X[:5])
print(probs.shape)
print(probs[0].sum())
print(model.classes_)
Part 4: Transformers - transform() and fit_transform()#
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
scaler.fit(X)
print(scaler.mean_.round(2))
X_scaled = scaler.transform(X)
print(X_scaled.mean(axis=0).round(2))
X_scaled_2 = StandardScaler().fit_transform(X)
print((X_scaled == X_scaled_2).all())
Part 5: Fitted Attributes - the Trailing Underscore Convention#
model = LogisticRegression(max_iter=200)
print(model.max_iter)
try:
print(model.coef_)
except AttributeError as e:
print(f'caught: {type(e).__name__}')
model.fit(X, y)
print(model.coef_.shape)
print(model.n_iter_)
Part 6: get_params() and set_params()#
model = LogisticRegression(max_iter=200, C=1.0)
params = model.get_params()
print(params['C'])
print(params['max_iter'])
model.set_params(C=0.5, max_iter=300)
print(model.C)
print(model.max_iter)
Part 7: Cloning Estimators with sklearn.base.clone#
from sklearn.base import clone
original = LogisticRegression(max_iter=200, C=2.0).fit(X, y)
copy = clone(original)
print(copy is original)
print(copy.C)
print(hasattr(copy, 'coef_'))
Part 8: score() - the Universal Evaluation Method#
from sklearn.linear_model import LinearRegression
from sklearn.datasets import load_diabetes
Xr, yr = load_diabetes(return_X_y=True)
clf = LogisticRegression(max_iter=200).fit(X, y)
print(clf.score(X, y))
reg = LinearRegression().fit(Xr, yr)
print(reg.score(Xr, yr))
Part 9: Checking Fitted State with check_is_fitted#
from sklearn.utils.validation import check_is_fitted
from sklearn.exceptions import NotFittedError
fresh_model = LogisticRegression(max_iter=200)
try:
check_is_fitted(fresh_model)
except NotFittedError as e:
print('caught: not yet fitted')
fresh_model.fit(X, y)
check_is_fitted(fresh_model)
print('fitted check passed silently')
Part 10: A Real Pattern - a Model-Agnostic Evaluation Helper#
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import train_test_split
def evaluate_any_model(model, X, y, test_size=0.3, random_state=42):
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=test_size, random_state=random_state)
model.fit(X_train, y_train)
return {'model': type(model).__name__, 'train_score': model.score(X_train, y_train), 'test_score': model.score(X_test, y_test)}
candidates = [LogisticRegression(max_iter=200), DecisionTreeClassifier(random_state=42), RandomForestClassifier(random_state=42)]
results = [evaluate_any_model(m, X, y) for m in candidates]
for r in results:
print(r)
Wrap-Up: What You Learned#
- Every scikit-learn estimator shares the identical fit/predict/transform/score interface, regardless of the algorithm underneath.
- fit always returns the estimator itself, enabling method chaining; constructors only set hyperparameters, never touch data.
- predict returns final labels/values; predict_proba returns per-class probabilities on classifiers that support it.
- transform applies what fit learned; fit_transform combines both steps, often more efficiently, on the same data.
- Fitted attributes always end in a trailing underscore, distinguishing learned state from constructor hyperparameters.
- get_params/set_params give a uniform way to read and update hyperparameters, exactly what GridSearchCV relies on internally.
- clone() produces a fresh, unfitted copy with identical hyperparameters, used internally by cross-validation and grid search.
- score() gives one universal evaluation number per estimator: accuracy for classifiers, R-squared for regressors, by default.
- check_is_fitted raises a clear, consistent error when predict/transform is called before fit.
- A real pattern: one small helper function can evaluate any estimator, thanks entirely to this shared, consistent API.
- That wraps up the estimator API. Next up: preprocessing with scalers, StandardScaler, MinMaxScaler, RobustScaler, and Normalizer.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



