Mathew K Analytics

Lesson 2 · Scikit-learn deep dive

Scikit-learn Tutorial #2: Preprocessing — Scalers

Video two of the eighteen-part series: scaling numeric features, and why it genuinely matters for many algorithms. StandardScaler, MinMaxScaler,…

⬇ Download notebookOpen in Colab ↗

📓 Full notebook

Download .ipynb

Scikit-learn Deep-Dive, Video 2: Preprocessing - Scalers#

  • Video two of the eighteen-part series: scaling numeric features, and why it genuinely matters for many algorithms.
  • StandardScaler, MinMaxScaler, RobustScaler, MaxAbsScaler, and Normalizer.
  • Let's get into it.

Part 1: Why Scaling Matters#

import numpy as np
from sklearn.datasets import load_wine
X, y = load_wine(return_X_y=True)
print(X[:, 0].min(), X[:, 0].max())
print(X[:, 9].min(), X[:, 9].max())
11.03 14.83
1.28 13.0

Part 2: StandardScaler - Zero Mean, Unit Variance#

from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
X_std = scaler.fit_transform(X)
print(X_std.mean(axis=0).round(2)[:3])
print(X_std.std(axis=0).round(2)[:3])
print(scaler.mean_[:3].round(2))
print(scaler.scale_[:3].round(2))
[ 0.  0. -0.]
[1. 1. 1.]
[13.    2.34  2.37]
[0.81 1.11 0.27]

Part 3: MinMaxScaler - Squeezing into a Fixed Range#

from sklearn.preprocessing import MinMaxScaler
mm_scaler = MinMaxScaler()
X_mm = mm_scaler.fit_transform(X)
print(X_mm.min(axis=0).round(2)[:3])
print(X_mm.max(axis=0).round(2)[:3])
custom_scaler = MinMaxScaler(feature_range=(-1, 1))
X_custom = custom_scaler.fit_transform(X)
print(X_custom.min(axis=0).round(2)[:3])
[0. 0. 0.]
[1. 1. 1.]
[-1. -1. -1.]

Part 4: RobustScaler - Resistant to Outliers#

from sklearn.preprocessing import RobustScaler, StandardScaler
data_with_outlier = np.array([[1], [2], [3], [4], [1000]], dtype=float)
std_result = StandardScaler().fit_transform(data_with_outlier)
robust_result = RobustScaler().fit_transform(data_with_outlier)
print(std_result.ravel().round(2))
print(robust_result.ravel().round(2))
[-0.5 -0.5 -0.5 -0.5  2. ]
[ -1.   -0.5   0.    0.5 498.5]

Part 5: Normalizer - Per-Sample, Not Per-Feature#

from sklearn.preprocessing import Normalizer
sample_data = np.array([[3, 4], [1, 1], [6, 8]], dtype=float)
normalizer = Normalizer(norm='l2')
normalized = normalizer.fit_transform(sample_data)
print(normalized.round(3))
print(np.linalg.norm(normalized, axis=1).round(3))
[[0.6   0.8  ]
 [0.707 0.707]
 [0.6   0.8  ]]
[1. 1. 1.]

Part 6: Fit on Train Only - Avoiding Data Leakage#

from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)
X_test_scaled = scaler.transform(X_test)
print(X_train_scaled.mean(axis=0).round(2)[:3])
print(X_test_scaled.mean(axis=0).round(2)[:3])
[ 0. -0.  0.]
[ 0.15 -0.2  -0.  ]

Part 7: inverse_transform - Undoing a Scaling#

scaler = StandardScaler().fit(X)
X_scaled = scaler.transform(X)
X_restored = scaler.inverse_transform(X_scaled)
print(np.allclose(X, X_restored))
print(X[0, :3].round(2))
print(X_restored[0, :3].round(2))
True
[14.23  1.71  2.43]
[14.23  1.71  2.43]

Part 8: Comparing Scalers Side by Side#

from sklearn.preprocessing import MaxAbsScaler
scalers = {
    'Standard': StandardScaler(),
    'MinMax': MinMaxScaler(),
    'Robust': RobustScaler(),
    'MaxAbs': MaxAbsScaler(),
}
for name, s in scalers.items():
    result = s.fit_transform(X)
    print(f'{name}: min={result.min():.2f}, max={result.max():.2f}')
Standard: min=-3.68, max=4.37
MinMax: min=0.00, max=1.00
Robust: min=-2.88, max=3.37
MaxAbs: min=0.07, max=1.00

Part 9: MaxAbsScaler - Preserving Sparsity#

from scipy import sparse
sparse_like = np.array([[0, 2, 0], [0, 0, 5], [1, 0, 0]], dtype=float)
mabs = MaxAbsScaler()
scaled_sparse = mabs.fit_transform(sparse_like)
print(scaled_sparse)
print((sparse_like == 0).sum() == (scaled_sparse == 0).sum())
[[0. 1. 0.]
 [0. 0. 1.]
 [1. 0. 0.]]
True

Part 10: A Real Pattern - Choosing the Right Scaler#

def recommend_scaler(has_outliers=False, needs_bounded_range=False, is_sparse=False):
    if is_sparse:
        return MaxAbsScaler()
    if has_outliers:
        return RobustScaler()
    if needs_bounded_range:
        return MinMaxScaler()
    return StandardScaler()
print(type(recommend_scaler(has_outliers=True)).__name__)
print(type(recommend_scaler(needs_bounded_range=True)).__name__)
print(type(recommend_scaler(is_sparse=True)).__name__)
print(type(recommend_scaler()).__name__)
RobustScaler
MinMaxScaler
MaxAbsScaler
StandardScaler

Wrap-Up: What You Learned#

  • Distance- and gradient-based algorithms are sensitive to raw feature scale; scaling puts features on comparable footing.
  • StandardScaler centers to zero mean, unit variance; it's the safe general default for roughly normal data.
  • MinMaxScaler squeezes into a fixed range, useful for algorithms expecting bounded input, but sensitive to outliers.
  • RobustScaler uses median and IQR instead of mean and standard deviation, genuinely resistant to outliers.
  • Normalizer rescales per sample (row-wise), not per feature (column-wise), for unit-norm, direction-focused data.
  • A scaler must be fit only on training data; fitting on the full dataset before splitting leaks test statistics into training.
  • Every scaler supports inverse_transform, mapping scaled values back to their original, meaningful units.
  • Because every scaler shares fit_transform, comparing several on the same data is just a simple loop.
  • MaxAbsScaler divides by the max absolute value, mapping to [-1, 1] without shifting; it never turns a zero nonzero, safe for sparse data.
  • A real pattern: a small decision function that recommends the right scaler based on the data's actual characteristics.
  • That wraps up scalers. Next up: encoders, OneHotEncoder, OrdinalEncoder, LabelEncoder, and LabelBinarizer.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.