Mathew K Analytics

Lesson 1 · Scikit-learn deep dive

Scikit-learn Tutorial #1: The Estimator API Explained

Video one of an eighteen-part series on scikit-learn's own toolkit, not the algorithms, the actual machinery that makes it a real production library. fit,…

⬇ Download notebookOpen in Colab ↗

📓 Full notebook

Download .ipynb

Scikit-learn Deep-Dive, Video 1: The Estimator API#

  • Video one of an eighteen-part series on scikit-learn's own toolkit, not the algorithms, the actual machinery that makes it a real production library.
  • fit, predict, transform, fitted attributes, get_params/set_params, cloning, and score.
  • Let's get into it.

Part 1: Why a Consistent API Matters#

from sklearn.linear_model import LogisticRegression
from sklearn.tree import DecisionTreeClassifier
from sklearn.datasets import load_iris
X, y = load_iris(return_X_y=True)
model_a = LogisticRegression(max_iter=200)
model_b = DecisionTreeClassifier()
model_a.fit(X, y)
model_b.fit(X, y)
print(model_a.predict(X[:3]))
print(model_b.predict(X[:3]))
[0 0 0]
[0 0 0]

Part 2: Estimators and fit() - the Universal Pattern#

model = LogisticRegression(max_iter=200)
print(hasattr(model, 'coef_'))
returned = model.fit(X, y)
print(returned is model)
print(hasattr(model, 'coef_'))
False
True
True

Part 3: Predictors - predict() and predict_proba()#

model = LogisticRegression(max_iter=200).fit(X, y)
preds = model.predict(X[:5])
print(preds)
probs = model.predict_proba(X[:5])
print(probs.shape)
print(probs[0].sum())
print(model.classes_)
[0 0 0 0 0]
(5, 3)
1.0
[0 1 2]

Part 4: Transformers - transform() and fit_transform()#

from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
scaler.fit(X)
print(scaler.mean_.round(2))
X_scaled = scaler.transform(X)
print(X_scaled.mean(axis=0).round(2))
X_scaled_2 = StandardScaler().fit_transform(X)
print((X_scaled == X_scaled_2).all())
[5.84 3.06 3.76 1.2 ]
[-0. -0. -0. -0.]
True

Part 5: Fitted Attributes - the Trailing Underscore Convention#

model = LogisticRegression(max_iter=200)
print(model.max_iter)
try:
    print(model.coef_)
except AttributeError as e:
    print(f'caught: {type(e).__name__}')
model.fit(X, y)
print(model.coef_.shape)
print(model.n_iter_)
200
caught: AttributeError
(3, 4)
[110]

Part 6: get_params() and set_params()#

model = LogisticRegression(max_iter=200, C=1.0)
params = model.get_params()
print(params['C'])
print(params['max_iter'])
model.set_params(C=0.5, max_iter=300)
print(model.C)
print(model.max_iter)
1.0
200
0.5
300

Part 7: Cloning Estimators with sklearn.base.clone#

from sklearn.base import clone
original = LogisticRegression(max_iter=200, C=2.0).fit(X, y)
copy = clone(original)
print(copy is original)
print(copy.C)
print(hasattr(copy, 'coef_'))
False
2.0
False

Part 8: score() - the Universal Evaluation Method#

from sklearn.linear_model import LinearRegression
from sklearn.datasets import load_diabetes
Xr, yr = load_diabetes(return_X_y=True)
clf = LogisticRegression(max_iter=200).fit(X, y)
print(clf.score(X, y))
reg = LinearRegression().fit(Xr, yr)
print(reg.score(Xr, yr))
0.9733333333333334
0.5177484222203499

Part 9: Checking Fitted State with check_is_fitted#

from sklearn.utils.validation import check_is_fitted
from sklearn.exceptions import NotFittedError
fresh_model = LogisticRegression(max_iter=200)
try:
    check_is_fitted(fresh_model)
except NotFittedError as e:
    print('caught: not yet fitted')
fresh_model.fit(X, y)
check_is_fitted(fresh_model)
print('fitted check passed silently')
caught: not yet fitted
fitted check passed silently

Part 10: A Real Pattern - a Model-Agnostic Evaluation Helper#

from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import train_test_split
def evaluate_any_model(model, X, y, test_size=0.3, random_state=42):
    X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=test_size, random_state=random_state)
    model.fit(X_train, y_train)
    return {'model': type(model).__name__, 'train_score': model.score(X_train, y_train), 'test_score': model.score(X_test, y_test)}
candidates = [LogisticRegression(max_iter=200), DecisionTreeClassifier(random_state=42), RandomForestClassifier(random_state=42)]
results = [evaluate_any_model(m, X, y) for m in candidates]
for r in results:
    print(r)
{'model': 'LogisticRegression', 'train_score': 0.9619047619047619, 'test_score': 1.0}
{'model': 'DecisionTreeClassifier', 'train_score': 1.0, 'test_score': 1.0}
{'model': 'RandomForestClassifier', 'train_score': 1.0, 'test_score': 1.0}

Wrap-Up: What You Learned#

  • Every scikit-learn estimator shares the identical fit/predict/transform/score interface, regardless of the algorithm underneath.
  • fit always returns the estimator itself, enabling method chaining; constructors only set hyperparameters, never touch data.
  • predict returns final labels/values; predict_proba returns per-class probabilities on classifiers that support it.
  • transform applies what fit learned; fit_transform combines both steps, often more efficiently, on the same data.
  • Fitted attributes always end in a trailing underscore, distinguishing learned state from constructor hyperparameters.
  • get_params/set_params give a uniform way to read and update hyperparameters, exactly what GridSearchCV relies on internally.
  • clone() produces a fresh, unfitted copy with identical hyperparameters, used internally by cross-validation and grid search.
  • score() gives one universal evaluation number per estimator: accuracy for classifiers, R-squared for regressors, by default.
  • check_is_fitted raises a clear, consistent error when predict/transform is called before fit.
  • A real pattern: one small helper function can evaluate any estimator, thanks entirely to this shared, consistent API.
  • That wraps up the estimator API. Next up: preprocessing with scalers, StandardScaler, MinMaxScaler, RobustScaler, and Normalizer.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.