Lesson 7 · Scikit-learn deep dive
Scikit-learn Tutorial #7: ColumnTransformer
Video seven of the eighteen-part series: applying different preprocessing to different columns. ColumnTransformer, makecolumntransformer, and mixed…
- CourseScikit-learn deep dive
- Lesson7 of 18
- Video15 min
- FormatJupyter notebook · 10 code cells
What you'll learn
- The Problem - Mixed Numeric and Categorical Columns
- Building a ColumnTransformer Explicitly
- remainder - passthrough vs drop
- makecolumntransformer - Auto-Naming
- ColumnTransformer Inside a Full Pipeline
- getfeaturenamesout on ColumnTransformer
- Selecting Columns by Dtype - makecolumnselector
- transformers - Inspecting Fitted Transformers
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbScikit-learn Deep-Dive, Video 7: ColumnTransformer#
- Video seven of the eighteen-part series: applying different preprocessing to different columns.
- ColumnTransformer, make_column_transformer, and mixed numeric/categorical data.
- Let's get into it.
Part 1: The Problem - Mixed Numeric and Categorical Columns#
import pandas as pd
import numpy as np
df = pd.DataFrame({
'age': [25, 32, 47, 51, 62, 29],
'income': [42000, 58000, 75000, 83000, 91000, 47000],
'city': ['NYC', 'LA', 'NYC', 'SF', 'LA', 'SF'],
'plan': ['basic', 'pro', 'pro', 'enterprise', 'basic', 'pro']
})
print(df.dtypes)
Part 2: Building a ColumnTransformer Explicitly#
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import StandardScaler, OneHotEncoder
numeric_cols = ['age', 'income']
categorical_cols = ['city', 'plan']
ct = ColumnTransformer([
('num', StandardScaler(), numeric_cols),
('cat', OneHotEncoder(handle_unknown='ignore'), categorical_cols)
])
transformed = ct.fit_transform(df)
print(transformed.shape)
Part 3: remainder - passthrough vs drop#
df2 = df.copy()
df2['id'] = [1, 2, 3, 4, 5, 6]
ct_drop = ColumnTransformer([('num', StandardScaler(), numeric_cols)])
ct_pass = ColumnTransformer([('num', StandardScaler(), numeric_cols)], remainder='passthrough')
print(ct_drop.fit_transform(df2).shape)
print(ct_pass.fit_transform(df2).shape)
Part 4: make_column_transformer - Auto-Naming#
from sklearn.compose import make_column_transformer
auto_ct = make_column_transformer(
(StandardScaler(), numeric_cols),
(OneHotEncoder(), categorical_cols)
)
print([name for name, _, _ in auto_ct.transformers])
auto_ct.fit(df)
print(auto_ct.transform(df).shape)
Part 5: ColumnTransformer Inside a Full Pipeline#
from sklearn.pipeline import Pipeline
from sklearn.linear_model import LogisticRegression
y = np.array([0, 1, 1, 1, 0, 1])
full_pipe = Pipeline([
('preprocess', ct),
('model', LogisticRegression(max_iter=1000))
])
full_pipe.fit(df, y)
print(round(full_pipe.score(df, y), 3))
Part 6: get_feature_names_out on ColumnTransformer#
ct_fitted = ColumnTransformer([
('num', StandardScaler(), numeric_cols),
('cat', OneHotEncoder(), categorical_cols)
]).fit(df)
print(ct_fitted.get_feature_names_out())
Part 7: Selecting Columns by Dtype - make_column_selector#
from sklearn.compose import make_column_selector
dynamic_ct = make_column_transformer(
(StandardScaler(), make_column_selector(dtype_include=np.number)),
(OneHotEncoder(), make_column_selector(dtype_include=object))
)
dynamic_ct.fit(df)
print(dynamic_ct.transform(df).shape)
Part 8: transformers_ - Inspecting Fitted Transformers#
for name, transformer, cols in dynamic_ct.transformers_:
print(name, cols)
Part 9: Cross-Validating the Full Mixed Pipeline Safely#
from sklearn.model_selection import cross_val_score
cv_scores = cross_val_score(full_pipe, df, y, cv=2)
print(cv_scores.round(3))
print(round(cv_scores.mean(), 3))
Part 10: A Real Pattern - a Reusable build_preprocessor Function#
def build_preprocessor(numeric, categorical):
return ColumnTransformer([
('num', StandardScaler(), numeric),
('cat', OneHotEncoder(handle_unknown='ignore'), categorical)
], remainder='drop')
preprocessor = build_preprocessor(numeric_cols, categorical_cols)
final_pipe = Pipeline([('preprocess', preprocessor), ('model', LogisticRegression(max_iter=1000))])
final_pipe.fit(df, y)
print(round(final_pipe.score(df, y), 3))
Wrap-Up: What You Learned#
- Real-world tables mix numeric and categorical columns, each needing different preprocessing.
- ColumnTransformer applies a different transformer to a different subset of columns, then concatenates the results.
- remainder controls unlisted columns: 'drop' (default) discards them, 'passthrough' keeps them unchanged.
- make_column_transformer auto-generates step names, mirroring make_pipeline.
- A fitted ColumnTransformer is itself a transformer, so it drops directly into a Pipeline as the first step.
- get_feature_names_out returns combined, correctly-prefixed names across every sub-transformer.
- make_column_selector picks columns dynamically by dtype or name pattern instead of a hardcoded list.
- transformers_ lists each fitted step alongside the exact columns it resolved to.
- Passing the whole ColumnTransformer-plus-model pipeline into cross_val_score keeps preprocessing leakage-safe on mixed data too.
- That wraps up ColumnTransformer. Next up: Model Selection, splitting strategies and cross-validation.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



