Mathew K Analytics

Lesson 7 · Data Mining

Understanding Data Transformation: Normalization, Encoding, and Feature Scaling in Machine Learning

Welcome to Week 3! This lesson is all about making data ready for analysis. You will learn about feature normalization, encoding, and scaling. These…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Week 3: Data Transformation in Data Mining#

Welcome to Week 3! This lesson is all about making data ready for analysis.

You will learn about feature normalization, encoding, and scaling.

These techniques are used every day in real-world data science and AI.

Ready to get started?

# Suppress warnings for a clean output
import warnings; warnings.filterwarnings('ignore')
import numpy as np
np.random.seed(42)
 
 

What is Data Transformation?#

Data transformation prepares raw data for machine learning.

Real-world data comes in different shapes, formats, and scales.

We use transformations to:

  • Clean and standardize inputs
  • Convert text or categories to numbers
  • Adjust values to a common range

Well-prepared data helps models understand and learn better.

# Data setup (Titanic Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
(891, 12)
   PassengerId  Survived  Pclass  \
0            1         0       3   
1            2         1       1   
2            3         1       3   

                                                Name     Sex   Age  SibSp  \
0                            Braund, Mr. Owen Harris    male  22.0      1   
1  Cumings, Mrs. John Bradley (Florence Briggs Th...  female  38.0      1   
2                             Heikkinen, Miss. Laina  female  26.0      0   

   Parch            Ticket     Fare Cabin Embarked  
0      0         A/5 21171   7.2500   NaN        S  
1      0          PC 17599  71.2833   C85        C  
2      0  STON/O2. 3101282   7.9250   NaN        S  

Why Transform Features?#

Machine learning algorithms work best when:

  • All features are numbers
  • Values have a similar scale
  • Missing values are filled or removed

Let us look at feature scaling, normalization, and encoding in action.

# Checking missing values
print(df.isnull().sum())
 
 
PassengerId      0
Survived         0
Pclass           0
Name             0
Sex              0
Age            177
SibSp            0
Parch            0
Ticket           0
Fare             0
Cabin          687
Embarked         2
dtype: int64
# Fill missing Age values with the median
median_age = df['Age'].median()
df['Age'] = df['Age'].fillna(median_age)
print(df['Age'].isnull().sum())
 
 
0

Normalization vs Standardization#

  • Normalization: Rescales features to a [0, 1] range.
  • Standardization: Changes features to have mean 0 and standard deviation 1.

Normalization is good for algorithms that compare distances, like k-NN.

Standardization is common for most machine learning algorithms.

# Selecting numerical columns for scaling
num_cols = ['Age', 'Fare']
print(df[num_cols].head())
 
 
    Age     Fare
0  22.0   7.2500
1  38.0  71.2833
2  26.0   7.9250
3  35.0  53.1000
4  35.0   8.0500
# Normalize Age and Fare between 0 and 1
from sklearn.preprocessing import MinMaxScaler
scaler = MinMaxScaler()
df_scaled = df.copy()
df_scaled[num_cols] = scaler.fit_transform(df[num_cols])
print(df_scaled[num_cols].head())
 
 
        Age      Fare
0  0.271174  0.014151
1  0.472229  0.139136
2  0.321438  0.015469
3  0.434531  0.103644
4  0.434531  0.015713
# Standardize Age and Fare
from sklearn.preprocessing import StandardScaler
std_scaler = StandardScaler()
df_standard = df.copy()
df_standard[num_cols] = std_scaler.fit_transform(df[num_cols])
print(df_standard[num_cols].head())
 
 
        Age      Fare
0 -0.565736 -0.502445
1  0.663861  0.786845
2 -0.258337 -0.488854
3  0.433312  0.420730
4  0.433312 -0.486337

What about text columns?#

Most algorithms need numbers, not text or categories.

We need to encode words into numbers.

Types of encoding:

  • Label encoding (assigns an integer for each category)
  • One-hot encoding (makes a column for each possible value)

Let us try encoding the Sex and Embarked columns.

# Label encoding the Sex column
from sklearn.preprocessing import LabelEncoder
le = LabelEncoder()
df_encoded = df.copy()
df_encoded['Sex'] = le.fit_transform(df['Sex'])
print(df_encoded[['Sex']].head(8))
 
 
   Sex
0    1
1    0
2    0
3    0
4    1
5    1
6    1
7    1
# One-hot encoding the Embarked column
df_ohe = pd.get_dummies(df, columns=['Embarked'], prefix='Emb')
print(df_ohe.filter(like='Emb').head(8))
 
 
   Emb_C  Emb_Q  Emb_S
0  False  False   True
1   True  False  False
2  False  False   True
3  False  False   True
4  False  False   True
5  False   True  False
6  False  False   True
7  False  False   True
# Combine scaling and encoding in one step
from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
num_features = ['Age','Fare']
cat_features = ['Sex','Embarked']
numeric_transformer = MinMaxScaler()
categorical_transformer = LabelEncoder()
# We need a helper for label encoding in pipelines
import numpy as np
from sklearn.base import BaseEstimator, TransformerMixin
class CustomLabelEncoder(BaseEstimator, TransformerMixin):
    def fit(self, X, y=None):
        self.le = LabelEncoder()
        self.le.fit(X)
        return self
    def transform(self, X):
        return self.le.transform(X).reshape(-1,1)
from sklearn.preprocessing import OneHotEncoder
preprocessor = ColumnTransformer(transformers=[
    ('num', MinMaxScaler(), num_features),
    ('cat', OneHotEncoder(), cat_features)
], remainder='drop')
X = df[num_features+cat_features].copy()
X_transformed = preprocessor.fit_transform(X)
print(X_transformed[:5])
 
 
[[0.27117366 0.01415106 0.         1.         0.         0.
  1.         0.        ]
 [0.4722292  0.13913574 1.         0.         1.         0.
  0.         0.        ]
 [0.32143755 0.01546857 1.         0.         0.         0.
  1.         0.        ]
 [0.43453129 0.1036443  1.         0.         0.         0.
  1.         0.        ]
 [0.43453129 0.01571255 0.         1.         0.         0.
  1.         0.        ]]

Why is scaling important?#

If we skip scaling, some features can dominate others.

For example, Fare might have much larger numbers than Age.

This can confuse models like k-NN or SVM, which use distances.

Always check that your features are on similar ranges before modeling!

# Mini-Project: Scale, Encode, and Split Data
from sklearn.model_selection import train_test_split
features = ['Pclass', 'Sex', 'Age', 'Fare', 'Embarked']
df_temp = df[features].copy()
# Fill missing Embarked
common_embarked = df_temp['Embarked'].mode()[0]
df_temp['Embarked'] = df_temp['Embarked'].fillna(common_embarked)
# One-hot encode Embarked, label encode Sex
df_temp['Sex'] = LabelEncoder().fit_transform(df_temp['Sex'])
df_temp = pd.get_dummies(df_temp, columns=['Embarked'], prefix='Emb')
# Normalize Age and Fare
scaler2 = MinMaxScaler()
df_temp[['Age','Fare']] = scaler2.fit_transform(df_temp[['Age','Fare']])
X = df_temp
y = df['Survived']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, random_state=42)
print('Train shape:', X_train.shape)
print('Test shape:', X_test.shape)
 
 
Train shape: (668, 7)
Test shape: (223, 7)
# Practice: Manually normalize a small list
data = [10, 20, 30, 40, 50]
min_val = min(data)
max_val = max(data)
data_norm = [(x - min_val) / (max_val - min_val) for x in data]
print('Original:', data)
print('Normalized:', data_norm)
 
 
Original: [10, 20, 30, 40, 50]
Normalized: [0.0, 0.25, 0.5, 0.75, 1.0]
# User input practice: Enter numbers to normalize
nums = input('Enter numbers separated by spaces: ')
nums = [float(x) for x in nums.split()]
min_n = min(nums)
max_n = max(nums)
normalized = [(x - min_n) / (max_n - min_n) for x in nums]
print('Normalized:', normalized)
 
 
Normalized: [0.0, 0.375, 0.625, 1.0]

Tips for Feature Scaling#

  • Always fit your scaler only on training data, not test data
  • Try different scalers if results look odd or slow to improve
  • Standardization works well for data with outliers
  • Normalize only the columns you need for your model

Remember, the right transformation can make or break your model!

# Challenge: Find all columns in Titanic data that are not numeric
non_numeric = df.select_dtypes(include=['object']).columns
print('Non-numeric columns:', list(non_numeric))
 
 
Non-numeric columns: ['Name', 'Sex', 'Ticket', 'Cabin', 'Embarked']

Recap: What did you learn today?#

  • Why we transform features before modeling
  • How to handle missing values
  • How to normalize and standardize data
  • Different types of encoding for categories
  • Putting it all together in a preprocessing workflow

Data transformation makes machine learning possible and accurate!

Thanks for Joining Week 3!#

Check out the description for more exercises and datasets.

If you want to practice more, comment below with your questions or ideas.

Subscribe for Week 4 and more data mining tutorials!

Happy coding and keep learning!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.