Lesson 8 · Data Mining
Understanding Principal Component Analysis (PCA) for Dimensionality Reduction in Data
Welcome! Today, we will learn how PCA (Principal Component Analysis) simplifies complex data, why it is useful, and how to use it with real datasets in…
- CourseData Mining
- Lesson8 of 31
- Video22 min
- FormatJupyter notebook · 15 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbWeek 3: Dimensionality Reduction with PCA#
Welcome! Today, we will learn how PCA (Principal Component Analysis) simplifies complex data, why it is useful, and how to use it with real datasets in Python.
By the end, you will know how to reduce unnecessary features and create clearer data visuals.
# To start, let us suppress warnings for a smoother learning experience.
import warnings
import numpy as np
np.random.seed(42)
warnings.filterwarnings("ignore")
What is Dimensionality Reduction?#
Machine learning data can have many features. Dimensionality reduction means lowering the number of these features while keeping the key information.
Why do we do it? Too many features can make models slow, confusing, or hard to visualize.
Imagine shopping online and having hundreds of filters. PCA helps us find which filters matter most.
# Data setup (Iris Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/uiuc-cse/data-fa14/gh-pages/data/iris.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
What is PCA?#
PCA stands for Principal Component Analysis.
It finds the directions that capture the most information in our data.
We use PCA to make big datasets simpler and to help visualize data in two or three dimensions.
# Let us look at our data types and check for missing values.
print(df.dtypes)
print(df.isnull().sum())
# PCA needs only the numeric features. Let us separate them now.
X = df.drop('species', axis=1)
print(X.head(3))
# Normalizing features: PCA works best when features are on the same scale.
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)
print(X_scaled[:3])
How does PCA work?#
PCA rotates your data to new directions, called principal components.
The first direction explains as much of your data as possible. The second direction explains the next-largest amount, and so on.
Think of it like finding the best angle to look at a busy street so you can see the most action.
# Let us apply PCA to our iris data. We will use only two principal components for visualization.
from sklearn.decomposition import PCA
pca = PCA(n_components=2)
X_pca = pca.fit_transform(X_scaled)
print('Shape after PCA:', X_pca.shape)
print(X_pca[:3])
# Let us add species labels back for plotting.
import numpy as np
pca_df = pd.DataFrame(X_pca, columns=['PC1', 'PC2'])
pca_df['species'] = df['species']
print(pca_df.head(3))
# Let us plot the PCA result to see if the species separate well.
import matplotlib.pyplot as plt
for s in pca_df['species'].unique():
plt.scatter(pca_df[pca_df['species']==s]['PC1'], pca_df[pca_df['species']==s]['PC2'], label=s)
plt.xlabel('PC1')
plt.ylabel('PC2')
plt.legend()
plt.title('Iris dataset visualized using PCA')
plt.show()
How many components should I keep?#
PCA lets you pick how many new features (components) to keep.
Usually, we keep as many as needed to explain most of the dataoften 90% or 95% of the variance.
# Let us see how much information each principal component keeps.
pca_full = PCA().fit(X_scaled)
explained_var = pca_full.explained_variance_ratio_
print('Variance explained by each component:', explained_var)
print('Total explained by first two:', explained_var[:2].sum())
# Let us try running PCA on the Titanic dataset for more practice.
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
titanic_df = pd.read_csv(url)
print(titanic_df.shape)
print(titanic_df.head(3))
# Let us preprocess Titanic: fill missing numbers and drop text clues.
num_cols = titanic_df.select_dtypes(include='number').columns
filled_df = titanic_df[num_cols].fillna(titanic_df[num_cols].mean())
print(filled_df.head(3))
# Now, let us scale and apply PCA to the Titanic numbers.
scaled_fill = scaler.fit_transform(filled_df)
titanic_pca = PCA(n_components=2)
titanic_pcs = titanic_pca.fit_transform(scaled_fill)
print('Shape:', titanic_pcs.shape)
print(titanic_pcs[:3])
# PRACTICE PROMPT: Try running this with three or more PCA components.
n = int(input("How many components would you like to use? (2-4): "))
titanic_pca_n = PCA(n_components=n)
titanic_pcs_n = titanic_pca_n.fit_transform(scaled_fill)
print(f'Titanic PCA with {n} components shape:', titanic_pcs_n.shape)
# Mini-project: Visualize Iris dataset in 3D using the three first PCA components.
pca_3d = PCA(n_components=3)
iris_pca3 = pca_3d.fit_transform(X_scaled)
from mpl_toolkits.mplot3d import Axes3D
fig = plt.figure(figsize=(8,6))
ax = fig.add_subplot(111, projection='3d')
colors = {'setosa':'r', 'versicolor':'g', 'virginica':'b'}
for s in pca_df['species'].unique():
label_mask = (df['species'] == s)
ax.scatter(iris_pca3[label_mask,0], iris_pca3[label_mask,1], iris_pca3[label_mask,2],
label=s, c=colors[s], s=50)
ax.set_xlabel('PC1')
ax.set_ylabel('PC2')
ax.set_zlabel('PC3')
plt.legend()
plt.title('Iris in 3D PCA space')
plt.show()
# Best Practices: Always scale before PCA and check explained variance.
# Try different numbers of components and look at the plot or explained variance.
Recap: Key Points about PCA#
PCA helps us explore high-dimensional data by keeping only what matters.
Always scale your features before PCA.
Use plots and explained variance to decide how many components to keep.
Try PCA on new datasets for practice.
Thank you for joining!#
If you liked this lesson, please subscribe to our channel and let us know your favorite PCA insight in the comments.
Practice: Try PCA with your favorite dataset.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



