Lesson 7 · Data Mining
Understanding Data Transformation: Normalization, Encoding, and Feature Scaling in Machine Learning
Welcome to Week 3! This lesson is all about making data ready for analysis. You will learn about feature normalization, encoding, and scaling. These…
- CourseData Mining
- Lesson7 of 31
- Video20 min
- FormatJupyter notebook · 14 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbWeek 3: Data Transformation in Data Mining#
Welcome to Week 3! This lesson is all about making data ready for analysis.
You will learn about feature normalization, encoding, and scaling.
These techniques are used every day in real-world data science and AI.
Ready to get started?
# Suppress warnings for a clean output
import warnings; warnings.filterwarnings('ignore')
import numpy as np
np.random.seed(42)
What is Data Transformation?#
Data transformation prepares raw data for machine learning.
Real-world data comes in different shapes, formats, and scales.
We use transformations to:
- Clean and standardize inputs
- Convert text or categories to numbers
- Adjust values to a common range
Well-prepared data helps models understand and learn better.
# Data setup (Titanic Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
Why Transform Features?#
Machine learning algorithms work best when:
- All features are numbers
- Values have a similar scale
- Missing values are filled or removed
Let us look at feature scaling, normalization, and encoding in action.
# Checking missing values
print(df.isnull().sum())
# Fill missing Age values with the median
median_age = df['Age'].median()
df['Age'] = df['Age'].fillna(median_age)
print(df['Age'].isnull().sum())
Normalization vs Standardization#
- Normalization: Rescales features to a [0, 1] range.
- Standardization: Changes features to have mean 0 and standard deviation 1.
Normalization is good for algorithms that compare distances, like k-NN.
Standardization is common for most machine learning algorithms.
# Selecting numerical columns for scaling
num_cols = ['Age', 'Fare']
print(df[num_cols].head())
# Normalize Age and Fare between 0 and 1
from sklearn.preprocessing import MinMaxScaler
scaler = MinMaxScaler()
df_scaled = df.copy()
df_scaled[num_cols] = scaler.fit_transform(df[num_cols])
print(df_scaled[num_cols].head())
# Standardize Age and Fare
from sklearn.preprocessing import StandardScaler
std_scaler = StandardScaler()
df_standard = df.copy()
df_standard[num_cols] = std_scaler.fit_transform(df[num_cols])
print(df_standard[num_cols].head())
What about text columns?#
Most algorithms need numbers, not text or categories.
We need to encode words into numbers.
Types of encoding:
- Label encoding (assigns an integer for each category)
- One-hot encoding (makes a column for each possible value)
Let us try encoding the Sex and Embarked columns.
# Label encoding the Sex column
from sklearn.preprocessing import LabelEncoder
le = LabelEncoder()
df_encoded = df.copy()
df_encoded['Sex'] = le.fit_transform(df['Sex'])
print(df_encoded[['Sex']].head(8))
# One-hot encoding the Embarked column
df_ohe = pd.get_dummies(df, columns=['Embarked'], prefix='Emb')
print(df_ohe.filter(like='Emb').head(8))
# Combine scaling and encoding in one step
from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
num_features = ['Age','Fare']
cat_features = ['Sex','Embarked']
numeric_transformer = MinMaxScaler()
categorical_transformer = LabelEncoder()
# We need a helper for label encoding in pipelines
import numpy as np
from sklearn.base import BaseEstimator, TransformerMixin
class CustomLabelEncoder(BaseEstimator, TransformerMixin):
def fit(self, X, y=None):
self.le = LabelEncoder()
self.le.fit(X)
return self
def transform(self, X):
return self.le.transform(X).reshape(-1,1)
from sklearn.preprocessing import OneHotEncoder
preprocessor = ColumnTransformer(transformers=[
('num', MinMaxScaler(), num_features),
('cat', OneHotEncoder(), cat_features)
], remainder='drop')
X = df[num_features+cat_features].copy()
X_transformed = preprocessor.fit_transform(X)
print(X_transformed[:5])
Why is scaling important?#
If we skip scaling, some features can dominate others.
For example, Fare might have much larger numbers than Age.
This can confuse models like k-NN or SVM, which use distances.
Always check that your features are on similar ranges before modeling!
# Mini-Project: Scale, Encode, and Split Data
from sklearn.model_selection import train_test_split
features = ['Pclass', 'Sex', 'Age', 'Fare', 'Embarked']
df_temp = df[features].copy()
# Fill missing Embarked
common_embarked = df_temp['Embarked'].mode()[0]
df_temp['Embarked'] = df_temp['Embarked'].fillna(common_embarked)
# One-hot encode Embarked, label encode Sex
df_temp['Sex'] = LabelEncoder().fit_transform(df_temp['Sex'])
df_temp = pd.get_dummies(df_temp, columns=['Embarked'], prefix='Emb')
# Normalize Age and Fare
scaler2 = MinMaxScaler()
df_temp[['Age','Fare']] = scaler2.fit_transform(df_temp[['Age','Fare']])
X = df_temp
y = df['Survived']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, random_state=42)
print('Train shape:', X_train.shape)
print('Test shape:', X_test.shape)
# Practice: Manually normalize a small list
data = [10, 20, 30, 40, 50]
min_val = min(data)
max_val = max(data)
data_norm = [(x - min_val) / (max_val - min_val) for x in data]
print('Original:', data)
print('Normalized:', data_norm)
# User input practice: Enter numbers to normalize
nums = input('Enter numbers separated by spaces: ')
nums = [float(x) for x in nums.split()]
min_n = min(nums)
max_n = max(nums)
normalized = [(x - min_n) / (max_n - min_n) for x in nums]
print('Normalized:', normalized)
Tips for Feature Scaling#
- Always fit your scaler only on training data, not test data
- Try different scalers if results look odd or slow to improve
- Standardization works well for data with outliers
- Normalize only the columns you need for your model
Remember, the right transformation can make or break your model!
# Challenge: Find all columns in Titanic data that are not numeric
non_numeric = df.select_dtypes(include=['object']).columns
print('Non-numeric columns:', list(non_numeric))
Recap: What did you learn today?#
- Why we transform features before modeling
- How to handle missing values
- How to normalize and standardize data
- Different types of encoding for categories
- Putting it all together in a preprocessing workflow
Data transformation makes machine learning possible and accurate!
Thanks for Joining Week 3!#
Check out the description for more exercises and datasets.
If you want to practice more, comment below with your questions or ideas.
Subscribe for Week 4 and more data mining tutorials!
Happy coding and keep learning!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



