Mathew K Analytics

Lesson 13 · Scikit-learn deep dive

Scikit-learn Tutorial #13: Feature Selection

Video thirteen of the eighteen-part series: keeping only the features genuinely worth having. VarianceThreshold, SelectKBest, RFE, and featureimportances.…

⬇ Download notebookOpen in Colab ↗

📓 Full notebook

Download .ipynb

Scikit-learn Deep-Dive, Video 13: Feature Selection#

  • Video thirteen of the eighteen-part series: keeping only the features genuinely worth having.
  • VarianceThreshold, SelectKBest, RFE, and feature_importances_.
  • Let's get into it.

Part 1: Why Select Features At All#

from sklearn.datasets import load_breast_cancer
X, y = load_breast_cancer(return_X_y=True)
print(X.shape)
(569, 30)

Part 2: VarianceThreshold - Removing Near-Constant Features#

import numpy as np
from sklearn.feature_selection import VarianceThreshold
X_with_constant = np.hstack([X, np.ones((X.shape[0], 1))])
vt = VarianceThreshold(threshold=0.0)
X_reduced = vt.fit_transform(X_with_constant)
print(X_with_constant.shape, X_reduced.shape)
(569, 31) (569, 30)

Part 3: SelectKBest with f_classif#

from sklearn.feature_selection import SelectKBest, f_classif
selector = SelectKBest(score_func=f_classif, k=10)
X_best = selector.fit_transform(X, y)
print(X_best.shape)
print(selector.scores_.round(1)[:5])
(569, 10)
[647.  118.1 697.2 573.1  83.7]

Part 4: SelectKBest with mutual_info_classif#

from sklearn.feature_selection import mutual_info_classif
mi_selector = SelectKBest(score_func=mutual_info_classif, k=10)
X_mi = mi_selector.fit_transform(X, y)
print(X_mi.shape)
(569, 10)

Part 5: get_support() - Which Features Were Kept#

feature_names = load_breast_cancer().feature_names
mask = selector.get_support()
kept_names = feature_names[mask]
print(list(kept_names))
[np.str_('mean radius'), np.str_('mean perimeter'), np.str_('mean area'), np.str_('mean concavity'), np.str_('mean concave points'), np.str_('worst radius'), np.str_('worst perimeter'), np.str_('worst area'), np.str_('worst concavity'), np.str_('worst concave points')]

Part 6: RFE - Recursive Feature Elimination#

from sklearn.feature_selection import RFE
from sklearn.linear_model import LogisticRegression
rfe = RFE(LogisticRegression(max_iter=5000), n_features_to_select=10)
rfe.fit(X, y)
print(rfe.support_.sum())
print(rfe.ranking_[:10])
10
[ 1  7 13 19  6  1  1  3  4 15]

Part 7: RFECV - Automatic k Selection#

from sklearn.feature_selection import RFECV
rfecv = RFECV(LogisticRegression(max_iter=5000), cv=5, min_features_to_select=5)
rfecv.fit(X, y)
print(rfecv.n_features_)
20

Part 8: feature_importances_ from Tree-Based Models#

from sklearn.ensemble import RandomForestClassifier
forest = RandomForestClassifier(random_state=42, n_estimators=100).fit(X, y)
importances = forest.feature_importances_
top5_idx = np.argsort(importances)[::-1][:5]
for idx in top5_idx:
    print(feature_names[idx], round(importances[idx], 3))
worst area 0.139
worst concave points 0.132
mean concave points 0.107
worst radius 0.083
worst perimeter 0.081

Part 9: SelectFromModel - Using a Model's Importances Directly#

from sklearn.feature_selection import SelectFromModel
sfm = SelectFromModel(forest, threshold='mean', prefit=True)
X_sfm = sfm.transform(X)
print(X_sfm.shape)
(569, 10)

Part 10: A Real Pattern - Feature Selection Inside a Pipeline#

from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import cross_val_score
select_pipe = Pipeline([
    ('scaler', StandardScaler()),
    ('select', SelectKBest(score_func=f_classif, k=10)),
    ('model', LogisticRegression(max_iter=5000))
])
scores = cross_val_score(select_pipe, X, y, cv=5)
print(round(scores.mean(), 3))
0.951

Wrap-Up: What You Learned#

  • Too many irrelevant or redundant features slow training, add noise, and hurt interpretability.
  • VarianceThreshold drops features whose variance falls below a threshold, catching near-constant columns.
  • SelectKBest keeps the top-k features by a scoring function; f_classif runs an ANOVA F-test.
  • mutual_info_classif captures non-linear relationships that a linear F-test can miss.
  • get_support() returns a boolean mask (or indices) showing exactly which features survived.
  • RFE repeatedly refits a model, dropping the weakest feature each round, until the target count remains.
  • RFECV cross-validates each candidate feature count, automatically choosing the best number.
  • Tree-based models expose feature_importances_ directly after fitting.
  • SelectFromModel wraps any fitted estimator, keeping features above a chosen importance threshold.
  • That wraps up feature selection. Next up: Text Data - Vectorization with CountVectorizer and TfidfVectorizer.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.