Lesson 60 · Python for Data Science
5 - Feature Scaling in Python
Feature scaling is a key step in preparing data for machine learning. It helps models learn better from your data. In this lesson, we will start simple and…
- CoursePython for Data Science
- Lesson60 of 38
- Video13 min
- FormatJupyter notebook · 22 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynb
Welcome to Feature Scaling in Python!#
Feature scaling is a key step in preparing data for machine learning. It helps models learn better from your data.
In this lesson, we will start simple and build up to small projects. You will learn what scaling is, why it matters, and how to use it in Python.
Why do we need feature scaling?#
Some models do not work well when features have very different scales. Feature scaling brings all the values into the same range.
This can make training faster and more accurate.
# Let us look at a simple example without scaling
heights = [150, 160, 170, 180]
weights = [50, 60, 70, 80]
print("Heights:", heights)
print("Weights:", weights)
# Now let us compute the minimum and maximum of heights
min_height = min(heights)
max_height = max(heights)
print("Min height:", min_height)
print("Max height:", max_height)
# And for weights
min_weight = min(weights)
max_weight = max(weights)
print("Min weight:", min_weight)
print("Max weight:", max_weight)
What is Min-Max Scaling?#
Min-Max scaling moves all values to a range between 0 and 1. This makes all the features equally important to the model. The formula is:
scaled_value = (value - min) / (max - min)
# Let us apply Min-Max scaling to heights using a loop
scaled_heights = []
for h in heights:
scaled = (h - min_height) / (max_height - min_height)
scaled_heights.append(scaled)
print("Scaled heights:", scaled_heights)
# Try scaling the weights too
scaled_weights = []
for w in weights:
scaled = (w - min_weight) / (max_weight - min_weight)
scaled_weights.append(scaled)
print("Scaled weights:", scaled_weights)
# Let's see Min-Max scaling in action with input
value = float(input("Enter a new height in cm: "))
scaled_value = (value - min_height) / (max_height - min_height)
print("Scaled value:", scaled_value)
What if a value is outside the original range?#
Min-Max scaling will give a value less than 0 or bigger than 1. This is called an outlier in scaling. Outliers can affect your models and results.
# Example: scaling a value outside the range
outside_value = 200
scaled_out = (outside_value - min_height) / (max_height - min_height)
print("Scaled result:", scaled_out)
# Standardization is another scaling method
import statistics
mean = statistics.mean(heights)
stdev = statistics.stdev(heights)
z_scores = []
for h in heights:
z = (h - mean) / stdev
z_scores.append(z)
print("Standardized heights (z-scores):", z_scores)
When should we scale features?#
Scaling is useful when:
- Features have very different scales
- Models use distances (like k-nearest neighbors or SVM)
Scaling is sometimes not needed for tree-based models like random forest.
# We can use NumPy for easier scaling
import numpy as np
heights_np = np.array(heights)
scaled_heights_np = (heights_np - heights_np.min()) / (heights_np.max() - heights_np.min())
print("NumPy scaled heights:", scaled_heights_np.tolist())
# Sklearn has built-in scalers for data science
from sklearn.preprocessing import MinMaxScaler
scaler = MinMaxScaler()
heights_arr = np.array(heights).reshape(-1, 1)
scaled = scaler.fit_transform(heights_arr)
print("Scaled heights with sklearn:", scaled.flatten().tolist())
# Scaling multiple features together
features = [[150, 50], [160, 60], [170, 70], [180, 80]]
scaler = MinMaxScaler()
scaled_features = scaler.fit_transform(features)
print("Scaled heights and weights together:")
print(scaled_features)
# Inverse transform changes scaled data back to the original
original = scaler.inverse_transform(scaled_features)
print("Back to original values:")
print(original)
# Let us try scaling with a real-world Iris dataset
from sklearn.datasets import load_iris
iris = load_iris()
data = iris.data
scaler = MinMaxScaler()
scaled_data = scaler.fit_transform(data)
print("First 5 rows of scaled Iris data:")
print(scaled_data[:5])
# Mini-project part 1: user scales their own two features
feat1 = float(input("First test score: "))
feat2 = float(input("Second test score: "))
all_scores = [feat1, feat2]
min_score = min(all_scores)
max_score = max(all_scores)
scaled_scores = [(s - min_score) / (max_score - min_score) for s in all_scores]
print("Original scores:", all_scores)
print("Scaled scores:", scaled_scores)
# Mini-project part 2: scaling ages for survey analysis
n = int(input("How many ages will you enter? "))
ages = []
for i in range(n):
age = float(input(f"Enter age #{i+1}: "))
ages.append(age)
min_age = min(ages)
max_age = max(ages)
scaled_ages = [(a - min_age) / (max_age - min_age) for a in ages]
print("Ages:", ages)
print("Scaled ages:", scaled_ages)
# Optimizing: building a reusable scaling function
def min_max_scale(data):
mn = min(data)
mx = max(data)
return [(x - mn) / (mx - mn) if mx != mn else 0 for x in data]
example = [4, 7, 10]
print("Scaled sample:", min_max_scale(example))
# Troubleshooting: handling empty lists and bad inputs
try:
test = []
print(min_max_scale(test))
except ValueError as e:
print("Error:", e)
try:
test = [5, 'bad', 9]
print(min_max_scale(test))
except Exception as e:
print("Error:", e)
# Extra tip: scaling with pandas DataFrame
import pandas as pd
df = pd.DataFrame({'a': [10, 20, 30], 'b': [1, 2, 3]})
df_scaled = (df - df.min()) / (df.max() - df.min())
print("Scaled DataFrame:")
print(df_scaled)
# Challenge: scale a list of temperatures between 0 and 1
temps = [23, 17, 30, 21, 25]
scaled_temps = min_max_scale(temps)
print("Temperatures:", temps)
print("Scaled temperatures:", scaled_temps)
Recap: What did we learn about feature scaling?#
- Why scaling matters for machine learning
- How to do Min-Max scaling and standardization
- How to use built-in Python, NumPy, pandas, and scikit-learn
- Mini-projects for practice
You can scale different kinds of data and prepare it for real-world analysis!
Thanks for learning Feature Scaling!#
If you enjoyed this, please:
- Like the video
- Comment with your questions
- Subscribe for more tutorials
- Share this lesson with friends
Happy Python coding!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



