Lesson 18 · Python For Machine Learning
Feature Scaling and Pipeline Construction in Python for Effective Machine Learning
In this lesson, we will learn how to prepare data for machine learning using feature scaling and pipelines. You will see why scaling matters and how…
- CoursePython For Machine Learning
- Lesson18 of 16
- Video12 min
- FormatJupyter notebook · 16 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynb
Welcome to Feature Scaling and Pipelines in Python#
In this lesson, we will learn how to prepare data for machine learning using feature scaling and pipelines.
You will see why scaling matters and how pipelines help keep code organized.
Let us get started!
# Always import the warnings library to filter out warnings
import warnings
warnings.filterwarnings("ignore") # This line hides most warning messages
What is Feature Scaling?#
Feature scaling means adjusting values in data so they fit into a certain range or distribution.
This is important because some machine learning models work better if all numbers are on the same scale.
For example: salaries and ages have very different values, but both affect predictions.
# Let us look at some numbers that need scaling
ages = [10, 18, 35, 45, 60]
salaries = [15000, 25000, 45000, 60000, 90000]
print("Ages:", ages)
print("Salaries:", salaries)
What is a Pipeline?#
A pipeline lets us link together all the data preparation steps and modeling into one simple workflow.
It is like an assembly line for your data.
Once set up, a pipeline makes your machine learning code neat and repeatable.
# Data setup: Let us load the Titanic dataset
import pandas as pd
url = "https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv"
df = pd.read_csv(url)
print("Data shape:", df.shape)
df.head()
# Let us check what columns we have and see the target variable distribution
print("Columns:", df.columns.tolist())
print(df['Survived'].value_counts())
# Let us focus on numerical columns for scaling
num_cols = ['Age', 'Fare']
df_num = df[num_cols]
df_num.head()
Why Scale Features?#
If numeric features are wildly different in size, some models pay more attention to big numbers.
Scaling makes sure every feature has a similar impact.
This is very useful for models like k-nearest neighbors or logistic regression.
# Standard scaling: zero mean, unit variance
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
df_scaled = scaler.fit_transform(df_num)
print(df_scaled[:5])
# Min-Max scaling example: scaling to range 0-1
from sklearn.preprocessing import MinMaxScaler
minmax_scaler = MinMaxScaler()
df_minmax = minmax_scaler.fit_transform(df_num)
print(df_minmax[:5])
What About Pipelines?#
Let us chain together preprocessing and a simple model using a pipeline.
This lets us handle all steps in one place and prevents mistakes.
# Separate our features (X) and label (y)
from sklearn.model_selection import train_test_split
X = df[['Age', 'Fare']].values
y = df['Survived'].values
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.3, random_state=42)
print('Train shape:', X_train.shape, 'Test shape:', X_test.shape)
# Build a pipeline: scale then fit logistic regression
from sklearn.pipeline import Pipeline
from sklearn.impute import SimpleImputer
from sklearn.linear_model import LogisticRegression
pipe = Pipeline([
("imputer", SimpleImputer(strategy="mean")),
("scaler", StandardScaler()),
("model", LogisticRegression(random_state=42))
])
pipe.fit(X_train, y_train)
score = pipe.score(X_test, y_test)
print("Test accuracy:", round(score, 2))
# Modify pipeline to use MinMaxScaler instead
pipe2 = Pipeline([
("imputer", SimpleImputer(strategy="mean")),
("scaler", MinMaxScaler()),
("model", LogisticRegression(random_state=42))
])
pipe2.fit(X_train, y_train)
score2 = pipe2.score(X_test, y_test)
print("Test accuracy with MinMaxScaler:", round(score2, 2))
Real-World Tip: Missing values#
Pipelines can also include steps to fill in missing values (imputation).
This keeps all your transformations together.
We often use SimpleImputer for this.
# Add a missing value imputer to pipeline
from sklearn.impute import SimpleImputer
pipe3 = Pipeline([
("imputer", SimpleImputer(strategy="mean")),
("scaler", StandardScaler()),
("model", LogisticRegression(random_state=42))
])
pipe3.fit(X_train, y_train)
score3 = pipe3.score(X_test, y_test)
print("Test accuracy with imputer:", round(score3, 2))
# Check if there are any missing values in Age or Fare
print("Missing Age:", df['Age'].isnull().sum())
print("Missing Fare:", df['Fare'].isnull().sum())
Best Practices: Train/Test Leakage#
Scaling must be fit only on training data, then applied to test data.
Pipelines make this automatic, so you will not leak test info to your model.
Always use pipelines or careful code to avoid data leakage.
# Practice: let us build a pipeline for a different model (KNeighborsClassifier)
from sklearn.neighbors import KNeighborsClassifier
pipe_knn = Pipeline([
("imputer", SimpleImputer(strategy="mean")),
("scaler", StandardScaler()),
("model", KNeighborsClassifier())
])
pipe_knn.fit(X_train, y_train)
score_knn = pipe_knn.score(X_test, y_test)
print("Test accuracy for k-NN:", round(score_knn, 2))
# Challenge: Try your own input! Enter age and fare to predict survival
age = float(input("Passenger Age: "))
fare = float(input("Ticket Fare: "))
prediction = pipe.predict([[age, fare]])
print("Predicted survival (0=No, 1=Yes):", prediction[0])
Recap: What Did We Learn?#
- Scaling puts features on equal footing for machine learning.
- Pipelines keep data steps together and avoid mistakes.
- Real-world data often needs both to work well.
Keep exploring and practicing your pipeline skills!
Thank you for joining this lesson on feature scaling and pipelines!
Give this video a thumbs up and subscribe for more easy Python tutorials.
See you next time happy coding!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



