Mathew K Analytics

Lesson 11 · ML algorithms deep dive

PCA & Dimensionality Reduction Explained | ML Algorithms #11

Video eleven of the 12-part series: compressing real high-dimensional data down to something you can actually plot. Real handwritten digit images, 1797 of…

What you'll learn

Datasets used in this lesson

Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.

📓 Full notebook

Download .ipynb

ML Algorithms Deep-Dive, Video 11: PCA and Dimensionality Reduction#

  • Video eleven of the 12-part series: compressing real high-dimensional data down to something you can actually plot.
  • Real handwritten digit images, 1797 of them, each a real 8 by 8 grid of pixels.
  • Let's get into it.

Before You Start#

  • Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
  • You'll need pandas, NumPy, Matplotlib, and scikit-learn.
  • Place digits.csv in the same folder as this notebook.
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
from sklearn.decomposition import PCA
from sklearn.neighbors import KNeighborsClassifier
from sklearn.model_selection import train_test_split, cross_val_score, StratifiedKFold
import time

digits = pd.read_csv('digits.csv')
X = digits.drop(columns=['target']).values
y = digits['target'].values
print(f'Real digit images: {len(digits)}')
print(f'Real pixel features per image: {X.shape[1]}')
Real digit images: 1797
Real pixel features per image: 64

Part 1: Seeing the Real Data#

fig, axes = plt.subplots(1, 5, figsize=(12, 3))
for i, ax in enumerate(axes):
    ax.imshow(X[i].reshape(8, 8), cmap='gray')
    ax.set_title(f'Real digit: {y[i]}')
    ax.axis('off')
plt.show()
No description has been provided for this image

Part 2: What PCA Actually Finds#

pca_full = PCA().fit(X)
explained = pca_full.explained_variance_ratio_
print(f'Real variance explained by component 1: {explained[0]:.4f}')
print(f'Real variance explained by component 2: {explained[1]:.4f}')
print(f'Real variance explained by the first 10 components combined: {explained[:10].sum():.4f}')
Real variance explained by component 1: 0.1489
Real variance explained by component 2: 0.1362
Real variance explained by the first 10 components combined: 0.7382

Part 3: How Many Components Are Actually Needed?#

cumulative_variance = np.cumsum(explained)
n_for_90 = np.argmax(cumulative_variance >= 0.90) + 1
n_for_95 = np.argmax(cumulative_variance >= 0.95) + 1
n_for_99 = np.argmax(cumulative_variance >= 0.99) + 1
print(f'Real components needed for 90% of variance: {n_for_90}')
print(f'Real components needed for 95% of variance: {n_for_95}')
print(f'Real components needed for 99% of variance: {n_for_99}')
Real components needed for 90% of variance: 21
Real components needed for 95% of variance: 29
Real components needed for 99% of variance: 41
plt.figure(figsize=(8, 5))
plt.plot(range(1, 65), cumulative_variance, color='steelblue')
plt.axhline(0.90, color='darkred', linestyle='--', label='90% variance')
plt.axvline(21, color='gray', linestyle=':')
plt.title('Real Cumulative Explained Variance')
plt.xlabel('Number of Components')
plt.ylabel('Cumulative Variance Explained')
plt.legend()
plt.show()
No description has been provided for this image

Part 4: Visualizing in Two Real Dimensions#

pca_2d = PCA(n_components=2).fit(X)
X_2d = pca_2d.transform(X)
print(f'Real variance captured by just these two components: {pca_2d.explained_variance_ratio_.sum():.4f}')
plt.figure(figsize=(9, 7))
scatter = plt.scatter(X_2d[:, 0], X_2d[:, 1], c=y, cmap='tab10', s=10)
plt.colorbar(scatter, label='Real Digit')
plt.title('Real Digits Projected onto Their First Two Principal Components')
plt.xlabel('Principal Component 1')
plt.ylabel('Principal Component 2')
plt.show()
Real variance captured by just these two components: 0.2851
No description has been provided for this image

Part 5: Reconstruction: What Gets Lost#

for n in [2, 10, 21, 30, 64]:
    pca_n = PCA(n_components=n).fit(X)
    X_reduced = pca_n.transform(X)
    X_reconstructed = pca_n.inverse_transform(X_reduced)
    mse = ((X - X_reconstructed) ** 2).mean()
    print(f'n_components={n}: real reconstruction MSE={mse:.4f}, real variance explained={pca_n.explained_variance_ratio_.sum():.4f}')
n_components=2: real reconstruction MSE=13.4210, real variance explained=0.2851
n_components=10: real reconstruction MSE=4.9143, real variance explained=0.7382
n_components=21: real reconstruction MSE=1.8173, real variance explained=0.9032
n_components=30: real reconstruction MSE=0.7681, real variance explained=0.9591
n_components=64: real reconstruction MSE=0.0000, real variance explained=1.0000
pca_recon = PCA(n_components=21).fit(X)
X_recon_21 = pca_recon.inverse_transform(pca_recon.transform(X))
fig, axes = plt.subplots(2, 5, figsize=(12, 5))
for i in range(5):
    axes[0, i].imshow(X[i].reshape(8, 8), cmap='gray')
    axes[0, i].set_title('Original')
    axes[0, i].axis('off')
    axes[1, i].imshow(X_recon_21[i].reshape(8, 8), cmap='gray')
    axes[1, i].set_title('21 Components')
    axes[1, i].axis('off')
plt.show()
No description has been provided for this image

Part 6: Does Compression Help a Real Classifier?#

X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, random_state=42, stratify=y)
skf = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
start_time = time.time()
raw_cv = cross_val_score(KNeighborsClassifier(n_neighbors=5), X_train, y_train, cv=skf, scoring='accuracy')
raw_time = time.time() - start_time
print(f'Real raw 64-dim CV accuracy: {raw_cv.mean():.4f}, real time: {raw_time:.3f}s')
Real raw 64-dim CV accuracy: 0.9851, real time: 2.272s
pca_21 = PCA(n_components=21).fit(X_train)
X_train_21 = pca_21.transform(X_train)
start_time = time.time()
pca_cv = cross_val_score(KNeighborsClassifier(n_neighbors=5), X_train_21, y_train, cv=skf, scoring='accuracy')
pca_time = time.time() - start_time
print(f'Real 21-dim PCA CV accuracy: {pca_cv.mean():.4f}, real time: {pca_time:.3f}s')
Real 21-dim PCA CV accuracy: 0.9837, real time: 0.134s

Part 7: Strengths, Weaknesses, and When to Use It#

  • Strength: genuinely speeds up distance- and memory-heavy real algorithms downstream, as part six showed directly.
  • Strength: makes real high-dimensional data visually inspectable, as part four showed.
  • Weakness: components are real linear combinations of original features, not directly interpretable the way a single real pixel or measurement is.
  • Weakness: always loses some genuine real information, the only question is how much, as part five made visible.
  • Use it as a real preprocessing step before a distance-based or otherwise expensive real algorithm on high-dimensional data, or purely for real exploratory visualization.

Wrap-Up: What You Learned#

  • PCA's real core idea: ranking directions in the data by how much genuine variance they capture.
  • Finding how many real components are actually needed, 21 of 64 for 90 percent of the real variance.
  • Visualizing an entire real 64-dimensional dataset in two dimensions, with an honest look at what that costs.
  • Reconstructing real compressed digits, and seeing they stay genuinely recognizable at a fraction of the original size.
  • A real, practical payoff: matching accuracy with a real classifier while running meaningfully faster.
  • Video twelve is the capstone, comparing every real algorithm from this series head to head on one final real dataset. Subscribe so it lands automatically see you there.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.