Mathew K Analytics

Lesson 8 · Data Mining

Understanding Principal Component Analysis (PCA) for Dimensionality Reduction in Data

Welcome! Today, we will learn how PCA (Principal Component Analysis) simplifies complex data, why it is useful, and how to use it with real datasets in…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Week 3: Dimensionality Reduction with PCA#

Welcome! Today, we will learn how PCA (Principal Component Analysis) simplifies complex data, why it is useful, and how to use it with real datasets in Python.

By the end, you will know how to reduce unnecessary features and create clearer data visuals.

# To start, let us suppress warnings for a smoother learning experience.
import warnings
import numpy as np
np.random.seed(42)
warnings.filterwarnings("ignore")

What is Dimensionality Reduction?#

Machine learning data can have many features. Dimensionality reduction means lowering the number of these features while keeping the key information.

Why do we do it? Too many features can make models slow, confusing, or hard to visualize.

Imagine shopping online and having hundreds of filters. PCA helps us find which filters matter most.

# Data setup (Iris Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/uiuc-cse/data-fa14/gh-pages/data/iris.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
(150, 5)
   sepal_length  sepal_width  petal_length  petal_width species
0           5.1          3.5           1.4          0.2  setosa
1           4.9          3.0           1.4          0.2  setosa
2           4.7          3.2           1.3          0.2  setosa

What is PCA?#

PCA stands for Principal Component Analysis.

It finds the directions that capture the most information in our data.

We use PCA to make big datasets simpler and to help visualize data in two or three dimensions.

# Let us look at our data types and check for missing values.
print(df.dtypes)
print(df.isnull().sum())
sepal_length    float64
sepal_width     float64
petal_length    float64
petal_width     float64
species          object
dtype: object
sepal_length    0
sepal_width     0
petal_length    0
petal_width     0
species         0
dtype: int64
# PCA needs only the numeric features. Let us separate them now.
X = df.drop('species', axis=1)
print(X.head(3))
   sepal_length  sepal_width  petal_length  petal_width
0           5.1          3.5           1.4          0.2
1           4.9          3.0           1.4          0.2
2           4.7          3.2           1.3          0.2
# Normalizing features: PCA works best when features are on the same scale.
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)
print(X_scaled[:3])
[[-0.90068117  1.03205722 -1.3412724  -1.31297673]
 [-1.14301691 -0.1249576  -1.3412724  -1.31297673]
 [-1.38535265  0.33784833 -1.39813811 -1.31297673]]

How does PCA work?#

PCA rotates your data to new directions, called principal components.

The first direction explains as much of your data as possible. The second direction explains the next-largest amount, and so on.

Think of it like finding the best angle to look at a busy street so you can see the most action.

# Let us apply PCA to our iris data. We will use only two principal components for visualization.
from sklearn.decomposition import PCA
pca = PCA(n_components=2)
X_pca = pca.fit_transform(X_scaled)
print('Shape after PCA:', X_pca.shape)
print(X_pca[:3])
Shape after PCA: (150, 2)
[[-2.26454173  0.5057039 ]
 [-2.0864255  -0.65540473]
 [-2.36795045 -0.31847731]]
# Let us add species labels back for plotting.
import numpy as np
pca_df = pd.DataFrame(X_pca, columns=['PC1', 'PC2'])
pca_df['species'] = df['species']
print(pca_df.head(3))
        PC1       PC2 species
0 -2.264542  0.505704  setosa
1 -2.086426 -0.655405  setosa
2 -2.367950 -0.318477  setosa
# Let us plot the PCA result to see if the species separate well.
import matplotlib.pyplot as plt
for s in pca_df['species'].unique():
    plt.scatter(pca_df[pca_df['species']==s]['PC1'], pca_df[pca_df['species']==s]['PC2'], label=s)
plt.xlabel('PC1')
plt.ylabel('PC2')
plt.legend()
plt.title('Iris dataset visualized using PCA')
plt.show()
No description has been provided for this image

How many components should I keep?#

PCA lets you pick how many new features (components) to keep.

Usually, we keep as many as needed to explain most of the dataoften 90% or 95% of the variance.

# Let us see how much information each principal component keeps.
pca_full = PCA().fit(X_scaled)
explained_var = pca_full.explained_variance_ratio_
print('Variance explained by each component:', explained_var)
print('Total explained by first two:', explained_var[:2].sum())
Variance explained by each component: [0.72770452 0.23030523 0.03683832 0.00515193]
Total explained by first two: 0.9580097536148195
# Let us try running PCA on the Titanic dataset for more practice.
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
titanic_df = pd.read_csv(url)
print(titanic_df.shape)
print(titanic_df.head(3))
(891, 12)
   PassengerId  Survived  Pclass  \
0            1         0       3   
1            2         1       1   
2            3         1       3   

                                                Name     Sex   Age  SibSp  \
0                            Braund, Mr. Owen Harris    male  22.0      1   
1  Cumings, Mrs. John Bradley (Florence Briggs Th...  female  38.0      1   
2                             Heikkinen, Miss. Laina  female  26.0      0   

   Parch            Ticket     Fare Cabin Embarked  
0      0         A/5 21171   7.2500   NaN        S  
1      0          PC 17599  71.2833   C85        C  
2      0  STON/O2. 3101282   7.9250   NaN        S  
# Let us preprocess Titanic: fill missing numbers and drop text clues.
num_cols = titanic_df.select_dtypes(include='number').columns
filled_df = titanic_df[num_cols].fillna(titanic_df[num_cols].mean())
print(filled_df.head(3))
   PassengerId  Survived  Pclass   Age  SibSp  Parch     Fare
0            1         0       3  22.0      1      0   7.2500
1            2         1       1  38.0      1      0  71.2833
2            3         1       3  26.0      0      0   7.9250
# Now, let us scale and apply PCA to the Titanic numbers.
scaled_fill = scaler.fit_transform(filled_df)
titanic_pca = PCA(n_components=2)
titanic_pcs = titanic_pca.fit_transform(scaled_fill)
print('Shape:', titanic_pcs.shape)
print(titanic_pcs[:3])
Shape: (891, 2)
[[ 1.4044005   0.43612559]
 [-2.01053121 -0.24285865]
 [ 0.49884943 -0.13774111]]
# PRACTICE PROMPT: Try running this with three or more PCA components.

n = int(input("How many components would you like to use? (2-4): "))
titanic_pca_n = PCA(n_components=n)
titanic_pcs_n = titanic_pca_n.fit_transform(scaled_fill)
print(f'Titanic PCA with {n} components shape:', titanic_pcs_n.shape)
 
 
Titanic PCA with 3 components shape: (891, 3)
# Mini-project: Visualize Iris dataset in 3D using the three first PCA components.
pca_3d = PCA(n_components=3)
iris_pca3 = pca_3d.fit_transform(X_scaled)
from mpl_toolkits.mplot3d import Axes3D
fig = plt.figure(figsize=(8,6))
ax = fig.add_subplot(111, projection='3d')
colors = {'setosa':'r', 'versicolor':'g', 'virginica':'b'}
for s in pca_df['species'].unique():
    label_mask = (df['species'] == s)
    ax.scatter(iris_pca3[label_mask,0], iris_pca3[label_mask,1], iris_pca3[label_mask,2],
               label=s, c=colors[s], s=50)
ax.set_xlabel('PC1')
ax.set_ylabel('PC2')
ax.set_zlabel('PC3')
plt.legend()
plt.title('Iris in 3D PCA space')
plt.show()
No description has been provided for this image
# Best Practices: Always scale before PCA and check explained variance.
# Try different numbers of components and look at the plot or explained variance.

Recap: Key Points about PCA#

  • PCA helps us explore high-dimensional data by keeping only what matters.

  • Always scale your features before PCA.

  • Use plots and explained variance to decide how many components to keep.

  • Try PCA on new datasets for practice.

Thank you for joining!#

If you liked this lesson, please subscribe to our channel and let us know your favorite PCA insight in the comments.

Practice: Try PCA with your favorite dataset.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.