Mathew K Analytics

Lesson 5 · Scikit-learn deep dive

Scikit-learn Tutorial #5: Feature Engineering

Video five of the eighteen-part series: building genuinely new, more useful features out of the ones you already have. PolynomialFeatures,…

⬇ Download notebookOpen in Colab ↗

📓 Full notebook

Download .ipynb

Scikit-learn Deep-Dive, Video 5: Feature Engineering#

  • Video five of the eighteen-part series: building genuinely new, more useful features out of the ones you already have.
  • PolynomialFeatures, FunctionTransformer, KBinsDiscretizer, and PowerTransformer.
  • Let's get into it.

Part 1: Why Engineer Features At All#

import numpy as np
from sklearn.linear_model import LinearRegression
X = np.linspace(-3, 3, 20).reshape(-1, 1)
y = (X.ravel() ** 2) + np.random.RandomState(0).normal(0, 0.5, 20)
plain_model = LinearRegression().fit(X, y)
print(round(plain_model.score(X, y), 3))
0.003

Part 2: PolynomialFeatures - degree#

from sklearn.preprocessing import PolynomialFeatures
poly = PolynomialFeatures(degree=2)
X_poly = poly.fit_transform(X)
print(X_poly[:3])
print(poly.get_feature_names_out(['x']))
[[ 1.         -3.          9.        ]
 [ 1.         -2.68421053  7.20498615]
 [ 1.         -2.36842105  5.60941828]]
['1' 'x' 'x^2']

Part 3: interaction_only and include_bias#

X2 = np.array([[1, 2], [3, 4], [5, 6]])
full_poly = PolynomialFeatures(degree=2).fit(X2)
interaction_poly = PolynomialFeatures(degree=2, interaction_only=True, include_bias=False).fit(X2)
print(full_poly.get_feature_names_out(['a', 'b']))
print(interaction_poly.get_feature_names_out(['a', 'b']))
['1' 'a' 'b' 'a^2' 'a b' 'b^2']
['a' 'b' 'a b']

Part 4: PolynomialFeatures with a Linear Model - Fitting a Curve#

poly2 = PolynomialFeatures(degree=2)
X_poly2 = poly2.fit_transform(X)
curved_model = LinearRegression().fit(X_poly2, y)
print(round(curved_model.score(X_poly2, y), 3))
print(curved_model.coef_.round(3))
0.983
[ 0.    -0.095  1.009]

Part 5: FunctionTransformer - Wrapping a Plain Function#

from sklearn.preprocessing import FunctionTransformer
skewed_data = np.array([[1], [2], [5], [10], [100], [1000]], dtype=float)
log_transformer = FunctionTransformer(np.log1p)
log_result = log_transformer.fit_transform(skewed_data)
print(log_result.ravel().round(2))
[0.69 1.1  1.79 2.4  4.62 6.91]

Part 6: FunctionTransformer with inverse_func#

log_transformer_inv = FunctionTransformer(np.log1p, inverse_func=np.expm1)
transformed = log_transformer_inv.fit_transform(skewed_data)
restored = log_transformer_inv.inverse_transform(transformed)
print(np.allclose(skewed_data, restored))
True

Part 7: FunctionTransformer with kw_args - Custom Logic#

def clip_and_scale(X, lower, upper):
    return np.clip(X, lower, upper) / upper
clipper = FunctionTransformer(clip_and_scale, kw_args={'lower': 0, 'upper': 100})
raw_values = np.array([[-10], [50], [150], [30]], dtype=float)
clipped_result = clipper.fit_transform(raw_values)
print(clipped_result.ravel())
[0.  0.5 1.  0.3]

Part 8: KBinsDiscretizer - Binning Continuous Features#

from sklearn.preprocessing import KBinsDiscretizer
ages = np.array([5, 17, 25, 40, 62, 80, 33, 19]).reshape(-1, 1)
binner = KBinsDiscretizer(n_bins=4, encode='ordinal', strategy='quantile', quantile_method='averaged_inverted_cdf')
binned = binner.fit_transform(ages)
print(binned.ravel())
print(binner.bin_edges_)
[0. 0. 1. 2. 3. 3. 2. 1.]
[array([ 5., 18., 29., 51., 80.])]

Part 9: PowerTransformer - Making Skewed Data More Gaussian#

from sklearn.preprocessing import PowerTransformer
skew_sample = np.random.RandomState(0).exponential(scale=2, size=(200, 1))
pt = PowerTransformer(method='yeo-johnson')
transformed_power = pt.fit_transform(skew_sample)
print(skew_sample.ravel()[:5].round(2))
print(transformed_power.ravel()[:5].round(2))
print(pt.lambdas_.round(3))
[1.59 2.51 1.85 1.57 1.1 ]
[ 0.15  0.67  0.32  0.14 -0.24]
[-0.328]

Part 10: A Real Pattern - a Reusable Feature Engineering Function#

def engineer_features(X, poly_degree=2):
    log_step = FunctionTransformer(np.log1p)
    poly_step = PolynomialFeatures(degree=poly_degree, include_bias=False)
    X_log = log_step.fit_transform(X)
    X_final = poly_step.fit_transform(X_log)
    return X_final, poly_step.get_feature_names_out()
sample = np.array([[1], [4], [9], [16]], dtype=float)
engineered, names = engineer_features(sample)
print(names)
print(engineered.round(3))
['x0' 'x0^2']
[[0.693 0.48 ]
 [1.609 2.59 ]
 [2.303 5.302]
 [2.833 8.027]]

Wrap-Up: What You Learned#

  • A linear model can only learn straight-line relationships in the raw features it's given; engineered features let it capture more.
  • PolynomialFeatures generates powers and combinations of features up to a given degree, letting a linear model fit curves.
  • interaction_only keeps only products of distinct features; include_bias controls whether a constant column of ones is added.
  • FunctionTransformer wraps any plain Python function as a real fit/transform-compatible estimator.
  • inverse_func gives a FunctionTransformer a real inverse_transform, for mapping predictions back to meaningful units.
  • kw_args passes extra keyword arguments through to the wrapped function, turning fully custom logic into a reusable transformer.
  • KBinsDiscretizer converts a continuous feature into a fixed number of discrete bins.
  • PowerTransformer learns an optimal transform (Yeo-Johnson by default) to make skewed data look more normally distributed.
  • A real pattern: a single reusable function combining a log transform and a polynomial expansion, ready for a pipeline.
  • That wraps up feature engineering. Next up: Pipelines, chaining preprocessing and models, and avoiding data leakage.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.