Mathew K Analytics

Lesson 7 · Scikit-learn deep dive

Scikit-learn Tutorial #7: ColumnTransformer

Video seven of the eighteen-part series: applying different preprocessing to different columns. ColumnTransformer, makecolumntransformer, and mixed…

⬇ Download notebookOpen in Colab ↗

📓 Full notebook

Download .ipynb

Scikit-learn Deep-Dive, Video 7: ColumnTransformer#

  • Video seven of the eighteen-part series: applying different preprocessing to different columns.
  • ColumnTransformer, make_column_transformer, and mixed numeric/categorical data.
  • Let's get into it.

Part 1: The Problem - Mixed Numeric and Categorical Columns#

import pandas as pd
import numpy as np
df = pd.DataFrame({
    'age': [25, 32, 47, 51, 62, 29],
    'income': [42000, 58000, 75000, 83000, 91000, 47000],
    'city': ['NYC', 'LA', 'NYC', 'SF', 'LA', 'SF'],
    'plan': ['basic', 'pro', 'pro', 'enterprise', 'basic', 'pro']
})
print(df.dtypes)
age        int64
income     int64
city      object
plan      object
dtype: object

Part 2: Building a ColumnTransformer Explicitly#

from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import StandardScaler, OneHotEncoder
numeric_cols = ['age', 'income']
categorical_cols = ['city', 'plan']
ct = ColumnTransformer([
    ('num', StandardScaler(), numeric_cols),
    ('cat', OneHotEncoder(handle_unknown='ignore'), categorical_cols)
])
transformed = ct.fit_transform(df)
print(transformed.shape)
(6, 8)

Part 3: remainder - passthrough vs drop#

df2 = df.copy()
df2['id'] = [1, 2, 3, 4, 5, 6]
ct_drop = ColumnTransformer([('num', StandardScaler(), numeric_cols)])
ct_pass = ColumnTransformer([('num', StandardScaler(), numeric_cols)], remainder='passthrough')
print(ct_drop.fit_transform(df2).shape)
print(ct_pass.fit_transform(df2).shape)
(6, 2)
(6, 5)

Part 4: make_column_transformer - Auto-Naming#

from sklearn.compose import make_column_transformer
auto_ct = make_column_transformer(
    (StandardScaler(), numeric_cols),
    (OneHotEncoder(), categorical_cols)
)
print([name for name, _, _ in auto_ct.transformers])
auto_ct.fit(df)
print(auto_ct.transform(df).shape)
['standardscaler', 'onehotencoder']
(6, 8)

Part 5: ColumnTransformer Inside a Full Pipeline#

from sklearn.pipeline import Pipeline
from sklearn.linear_model import LogisticRegression
y = np.array([0, 1, 1, 1, 0, 1])
full_pipe = Pipeline([
    ('preprocess', ct),
    ('model', LogisticRegression(max_iter=1000))
])
full_pipe.fit(df, y)
print(round(full_pipe.score(df, y), 3))
1.0

Part 6: get_feature_names_out on ColumnTransformer#

ct_fitted = ColumnTransformer([
    ('num', StandardScaler(), numeric_cols),
    ('cat', OneHotEncoder(), categorical_cols)
]).fit(df)
print(ct_fitted.get_feature_names_out())
['num__age' 'num__income' 'cat__city_LA' 'cat__city_NYC' 'cat__city_SF'
 'cat__plan_basic' 'cat__plan_enterprise' 'cat__plan_pro']

Part 7: Selecting Columns by Dtype - make_column_selector#

from sklearn.compose import make_column_selector
dynamic_ct = make_column_transformer(
    (StandardScaler(), make_column_selector(dtype_include=np.number)),
    (OneHotEncoder(), make_column_selector(dtype_include=object))
)
dynamic_ct.fit(df)
print(dynamic_ct.transform(df).shape)
(6, 8)

Part 8: transformers_ - Inspecting Fitted Transformers#

for name, transformer, cols in dynamic_ct.transformers_:
    print(name, cols)
standardscaler ['age', 'income']
onehotencoder ['city', 'plan']

Part 9: Cross-Validating the Full Mixed Pipeline Safely#

from sklearn.model_selection import cross_val_score
cv_scores = cross_val_score(full_pipe, df, y, cv=2)
print(cv_scores.round(3))
print(round(cv_scores.mean(), 3))
[0.667 0.667]
0.667

Part 10: A Real Pattern - a Reusable build_preprocessor Function#

def build_preprocessor(numeric, categorical):
    return ColumnTransformer([
        ('num', StandardScaler(), numeric),
        ('cat', OneHotEncoder(handle_unknown='ignore'), categorical)
    ], remainder='drop')
preprocessor = build_preprocessor(numeric_cols, categorical_cols)
final_pipe = Pipeline([('preprocess', preprocessor), ('model', LogisticRegression(max_iter=1000))])
final_pipe.fit(df, y)
print(round(final_pipe.score(df, y), 3))
1.0

Wrap-Up: What You Learned#

  • Real-world tables mix numeric and categorical columns, each needing different preprocessing.
  • ColumnTransformer applies a different transformer to a different subset of columns, then concatenates the results.
  • remainder controls unlisted columns: 'drop' (default) discards them, 'passthrough' keeps them unchanged.
  • make_column_transformer auto-generates step names, mirroring make_pipeline.
  • A fitted ColumnTransformer is itself a transformer, so it drops directly into a Pipeline as the first step.
  • get_feature_names_out returns combined, correctly-prefixed names across every sub-transformer.
  • make_column_selector picks columns dynamically by dtype or name pattern instead of a hardcoded list.
  • transformers_ lists each fitted step alongside the exact columns it resolved to.
  • Passing the whole ColumnTransformer-plus-model pipeline into cross_val_score keeps preprocessing leakage-safe on mixed data too.
  • That wraps up ColumnTransformer. Next up: Model Selection, splitting strategies and cross-validation.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.