Lesson 11 · ML algorithms deep dive
PCA & Dimensionality Reduction Explained | ML Algorithms #11
Video eleven of the 12-part series: compressing real high-dimensional data down to something you can actually plot. Real handwritten digit images, 1797 of…
- CourseML algorithms deep dive
- Lesson11 of 12
- Video15 min
- FormatJupyter notebook · 10 code cells
- Data1 dataset
What you'll learn
Datasets used in this lesson
Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.
- digits.csv483.8 KB
📓 Full notebook
Download .ipynbML Algorithms Deep-Dive, Video 11: PCA and Dimensionality Reduction#
- Video eleven of the 12-part series: compressing real high-dimensional data down to something you can actually plot.
- Real handwritten digit images, 1797 of them, each a real 8 by 8 grid of pixels.
- Let's get into it.
Before You Start#
- Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
- You'll need pandas, NumPy, Matplotlib, and scikit-learn.
- Place digits.csv in the same folder as this notebook.
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
from sklearn.decomposition import PCA
from sklearn.neighbors import KNeighborsClassifier
from sklearn.model_selection import train_test_split, cross_val_score, StratifiedKFold
import time
digits = pd.read_csv('digits.csv')
X = digits.drop(columns=['target']).values
y = digits['target'].values
print(f'Real digit images: {len(digits)}')
print(f'Real pixel features per image: {X.shape[1]}')
Part 1: Seeing the Real Data#
fig, axes = plt.subplots(1, 5, figsize=(12, 3))
for i, ax in enumerate(axes):
ax.imshow(X[i].reshape(8, 8), cmap='gray')
ax.set_title(f'Real digit: {y[i]}')
ax.axis('off')
plt.show()
Part 2: What PCA Actually Finds#
pca_full = PCA().fit(X)
explained = pca_full.explained_variance_ratio_
print(f'Real variance explained by component 1: {explained[0]:.4f}')
print(f'Real variance explained by component 2: {explained[1]:.4f}')
print(f'Real variance explained by the first 10 components combined: {explained[:10].sum():.4f}')
Part 3: How Many Components Are Actually Needed?#
cumulative_variance = np.cumsum(explained)
n_for_90 = np.argmax(cumulative_variance >= 0.90) + 1
n_for_95 = np.argmax(cumulative_variance >= 0.95) + 1
n_for_99 = np.argmax(cumulative_variance >= 0.99) + 1
print(f'Real components needed for 90% of variance: {n_for_90}')
print(f'Real components needed for 95% of variance: {n_for_95}')
print(f'Real components needed for 99% of variance: {n_for_99}')
plt.figure(figsize=(8, 5))
plt.plot(range(1, 65), cumulative_variance, color='steelblue')
plt.axhline(0.90, color='darkred', linestyle='--', label='90% variance')
plt.axvline(21, color='gray', linestyle=':')
plt.title('Real Cumulative Explained Variance')
plt.xlabel('Number of Components')
plt.ylabel('Cumulative Variance Explained')
plt.legend()
plt.show()
Part 4: Visualizing in Two Real Dimensions#
pca_2d = PCA(n_components=2).fit(X)
X_2d = pca_2d.transform(X)
print(f'Real variance captured by just these two components: {pca_2d.explained_variance_ratio_.sum():.4f}')
plt.figure(figsize=(9, 7))
scatter = plt.scatter(X_2d[:, 0], X_2d[:, 1], c=y, cmap='tab10', s=10)
plt.colorbar(scatter, label='Real Digit')
plt.title('Real Digits Projected onto Their First Two Principal Components')
plt.xlabel('Principal Component 1')
plt.ylabel('Principal Component 2')
plt.show()
Part 5: Reconstruction: What Gets Lost#
for n in [2, 10, 21, 30, 64]:
pca_n = PCA(n_components=n).fit(X)
X_reduced = pca_n.transform(X)
X_reconstructed = pca_n.inverse_transform(X_reduced)
mse = ((X - X_reconstructed) ** 2).mean()
print(f'n_components={n}: real reconstruction MSE={mse:.4f}, real variance explained={pca_n.explained_variance_ratio_.sum():.4f}')
pca_recon = PCA(n_components=21).fit(X)
X_recon_21 = pca_recon.inverse_transform(pca_recon.transform(X))
fig, axes = plt.subplots(2, 5, figsize=(12, 5))
for i in range(5):
axes[0, i].imshow(X[i].reshape(8, 8), cmap='gray')
axes[0, i].set_title('Original')
axes[0, i].axis('off')
axes[1, i].imshow(X_recon_21[i].reshape(8, 8), cmap='gray')
axes[1, i].set_title('21 Components')
axes[1, i].axis('off')
plt.show()
Part 6: Does Compression Help a Real Classifier?#
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, random_state=42, stratify=y)
skf = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
start_time = time.time()
raw_cv = cross_val_score(KNeighborsClassifier(n_neighbors=5), X_train, y_train, cv=skf, scoring='accuracy')
raw_time = time.time() - start_time
print(f'Real raw 64-dim CV accuracy: {raw_cv.mean():.4f}, real time: {raw_time:.3f}s')
pca_21 = PCA(n_components=21).fit(X_train)
X_train_21 = pca_21.transform(X_train)
start_time = time.time()
pca_cv = cross_val_score(KNeighborsClassifier(n_neighbors=5), X_train_21, y_train, cv=skf, scoring='accuracy')
pca_time = time.time() - start_time
print(f'Real 21-dim PCA CV accuracy: {pca_cv.mean():.4f}, real time: {pca_time:.3f}s')
Part 7: Strengths, Weaknesses, and When to Use It#
- Strength: genuinely speeds up distance- and memory-heavy real algorithms downstream, as part six showed directly.
- Strength: makes real high-dimensional data visually inspectable, as part four showed.
- Weakness: components are real linear combinations of original features, not directly interpretable the way a single real pixel or measurement is.
- Weakness: always loses some genuine real information, the only question is how much, as part five made visible.
- Use it as a real preprocessing step before a distance-based or otherwise expensive real algorithm on high-dimensional data, or purely for real exploratory visualization.
Wrap-Up: What You Learned#
- PCA's real core idea: ranking directions in the data by how much genuine variance they capture.
- Finding how many real components are actually needed, 21 of 64 for 90 percent of the real variance.
- Visualizing an entire real 64-dimensional dataset in two dimensions, with an honest look at what that costs.
- Reconstructing real compressed digits, and seeing they stay genuinely recognizable at a fraction of the original size.
- A real, practical payoff: matching accuracy with a real classifier while running meaningfully faster.
- Video twelve is the capstone, comparing every real algorithm from this series head to head on one final real dataset. Subscribe so it lands automatically see you there.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



