Mathew K Analytics

Lesson 4 · Scikit-learn deep dive

Scikit-learn Tutorial #4: Handling Missing Data

Video four of the eighteen-part series: real strategies for filling in genuine gaps in a dataset. SimpleImputer, MissingIndicator, KNNImputer, and…

⬇ Download notebookOpen in Colab ↗

📓 Full notebook

Download .ipynb

Scikit-learn Deep-Dive, Video 4: Handling Missing Data#

  • Video four of the eighteen-part series: real strategies for filling in genuine gaps in a dataset.
  • SimpleImputer, MissingIndicator, KNNImputer, and IterativeImputer.
  • Let's get into it.

Part 1: Why Missing Data Needs Explicit Handling#

import numpy as np
from sklearn.linear_model import LinearRegression
X = np.array([[1.0, 2.0], [3.0, np.nan], [5.0, 6.0], [np.nan, 8.0]])
y = np.array([1, 2, 3, 4])
try:
    LinearRegression().fit(X, y)
except ValueError as e:
    print('caught: NaN rejected')
caught: NaN rejected

Part 2: SimpleImputer - mean, median, most_frequent, constant#

from sklearn.impute import SimpleImputer
mean_imputer = SimpleImputer(strategy='mean')
X_mean = mean_imputer.fit_transform(X)
print(X_mean)
print(mean_imputer.statistics_)
median_imputer = SimpleImputer(strategy='median')
X_median = median_imputer.fit_transform(X)
print(X_median)
[[1.         2.        ]
 [3.         5.33333333]
 [5.         6.        ]
 [3.         8.        ]]
[3.         5.33333333]
[[1. 2.]
 [3. 6.]
 [5. 6.]
 [3. 8.]]

Part 3: SimpleImputer on Categorical Data#

cat_data = np.array([['red'], ['blue'], [None], ['red'], ['red']], dtype=object)
freq_imputer = SimpleImputer(strategy='most_frequent', missing_values=None)
cat_filled = freq_imputer.fit_transform(cat_data)
print(cat_filled.ravel())
const_imputer = SimpleImputer(strategy='constant', fill_value='unknown', missing_values=None)
cat_const = const_imputer.fit_transform(cat_data)
print(cat_const.ravel())
['red' 'blue' 'red' 'red' 'red']
['red' 'blue' 'unknown' 'red' 'red']

Part 4: MissingIndicator - Tracking Where Data Was Missing#

from sklearn.impute import MissingIndicator
indicator = MissingIndicator()
missing_mask = indicator.fit_transform(X)
print(missing_mask)
print(indicator.features_)
[[False False]
 [False  True]
 [False False]
 [ True False]]
[0 1]

Part 5: add_indicator - Filling and Tracking in One Step#

combined_imputer = SimpleImputer(strategy='mean', add_indicator=True)
X_combined = combined_imputer.fit_transform(X)
print(X_combined.shape)
print(X_combined)
(4, 4)
[[1.         2.         0.         0.        ]
 [3.         5.33333333 0.         1.        ]
 [5.         6.         0.         0.        ]
 [3.         8.         1.         0.        ]]

Part 6: KNNImputer - Filling from Similar Rows#

from sklearn.impute import KNNImputer
X_bigger = np.array([[1, 2], [3, np.nan], [5, 6], [7, 8], [np.nan, 10], [2, 3]])
knn_imputer = KNNImputer(n_neighbors=2)
X_knn = knn_imputer.fit_transform(X_bigger)
print(X_knn)
[[ 1.   2. ]
 [ 3.   2.5]
 [ 5.   6. ]
 [ 7.   8. ]
 [ 6.  10. ]
 [ 2.   3. ]]

Part 7: IterativeImputer - Modeling Each Feature from the Others#

from sklearn.experimental import enable_iterative_imputer
from sklearn.impute import IterativeImputer
iter_imputer = IterativeImputer(random_state=42, max_iter=10)
X_iter = iter_imputer.fit_transform(X_bigger)
print(X_iter.round(2))
[[ 1.  2.]
 [ 3.  4.]
 [ 5.  6.]
 [ 7.  8.]
 [ 9. 10.]
 [ 2.  3.]]

Part 8: Fit on Train Only - the Identical Leakage Rule#

from sklearn.model_selection import train_test_split
X_train, X_test = train_test_split(X_bigger, test_size=0.3, random_state=42)
imputer = SimpleImputer(strategy='mean')
X_train_filled = imputer.fit_transform(X_train)
X_test_filled = imputer.transform(X_test)
print(imputer.statistics_)
print(X_test_filled)
[4.66666667 6.75      ]
[[1.   2.  ]
 [3.   6.75]]

Part 9: Comparing Imputation Strategies Side by Side#

strategies = {
    'mean': SimpleImputer(strategy='mean'),
    'median': SimpleImputer(strategy='median'),
    'knn': KNNImputer(n_neighbors=2),
}
for name, imp in strategies.items():
    filled = imp.fit_transform(X_bigger)
    print(f'{name}: row 1 col 1 filled as {filled[1, 1]:.2f}')
mean: row 1 col 1 filled as 5.80
median: row 1 col 1 filled as 6.00
knn: row 1 col 1 filled as 2.50

Part 10: A Real Pattern - a Missing-Data Strategy Helper#

def recommend_imputer(n_rows, missingness_may_be_meaningful=False, is_categorical=False):
    if is_categorical:
        base = SimpleImputer(strategy='most_frequent')
    elif n_rows < 1000:
        base = KNNImputer(n_neighbors=5)
    else:
        base = SimpleImputer(strategy='median')
    return base
print(type(recommend_imputer(500)).__name__)
print(type(recommend_imputer(50000)).__name__)
print(type(recommend_imputer(500, is_categorical=True)).__name__)
KNNImputer
SimpleImputer
SimpleImputer

Wrap-Up: What You Learned#

  • Most estimators reject NaN outright; missing data must be filled or explicitly tracked before it ever reaches fit.
  • SimpleImputer fills gaps with one per-feature statistic: mean/median for numeric, most_frequent/constant for any data type.
  • MissingIndicator produces a binary column marking exactly where every gap originally was, since missingness can itself be predictive.
  • SimpleImputer's add_indicator=True combines filling and tracking into a single transform call.
  • KNNImputer fills gaps from the average of the k nearest rows, often more accurate than one global statistic.
  • IterativeImputer models each feature with gaps as a regression target from the other features; still experimental, needs an explicit opt-in import.
  • Imputers follow the identical leakage rule as scalers: fit only on training data, transform (never fit_transform) on test data.
  • Because every imputer shares fit_transform, comparing several strategies on the same data is just a simple loop.
  • A real pattern: a small decision function recommending an imputation strategy based on data size, type, and whether missingness itself is meaningful.
  • That wraps up handling missing data. Next up: feature engineering with PolynomialFeatures and FunctionTransformer.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.