Lesson 68 · Python Fundamentals
Understanding Principal Component Analysis (PCA) in Python: A Comprehensive Tutorial
Principal Component Analysis, or PCA, is a popular way to understand and simplify complex datasets. Let us discover how PCA helps us find patterns, reduce…
- CoursePython Fundamentals
- Lesson68 of 22
- Video13 min
- FormatJupyter notebook · 15 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynb
Welcome to Principal Component Analysis (PCA) in Python#
Principal Component Analysis, or PCA, is a popular way to understand and simplify complex datasets. Let us discover how PCA helps us find patterns, reduce dimensions, and visualize data. We will use real examples, including the famous Iris flower dataset.
What is PCA?#
PCA finds patterns in data by looking for the directions where the data varies the most. It lets us make a big, complicated table of numbers into a small and simple oneso humans and computers can see important shapes.
# Let us set up Python to hide unnecessary warnings.
import warnings
warnings.filterwarnings('ignore')
# Data setup: Let us load and preview the Iris dataset.
from sklearn.datasets import load_iris
import pandas as pd
iris = load_iris()
X = pd.DataFrame(iris.data, columns=iris.feature_names)
y = pd.Series(iris.target, name='species')
print('Shape:', X.shape)
print('First 5 rows:')
print(X.head())
# Let us check how many samples are in each species group.
print('Species counts:')
print(y.value_counts())
# Let us view basic statistics for the features.
print(X.describe())
# Standardize the data to mean 0 and variance 1.
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)
print('Mean of each column after scaling:')
print(X_scaled.mean(axis=0))
Starting PCA#
Let us start finding the important directions in our data. We will use scikit-learn, a well-known Python library, to do the math for us.
# Fit a PCA model and look at explained variance.
from sklearn.decomposition import PCA
pca = PCA()
pca.fit(X_scaled)
print('Explained variance ratio:', pca.explained_variance_ratio_)
# Let us plot the cumulative variance explained.
import matplotlib.pyplot as plt
plt.plot(range(1, 1 + len(pca.explained_variance_ratio_)), pca.explained_variance_ratio_.cumsum(), marker='o')
plt.xlabel('Number of components')
plt.ylabel('Cumulative explained variance')
plt.title('Variance explained by PCA components')
plt.show()
# Transform data using two principal components for visualization.
pca2 = PCA(n_components=2)
X_pca2 = pca2.fit_transform(X_scaled)
print('Shape after PCA:', X_pca2.shape)
print('First few PCA-transformed rows:')
print(X_pca2[:5])
# Visualize the PCA result with a scatter plot.
plt.figure(figsize=(6,4))
for lab, col in zip(range(3), ['red', 'green', 'blue']):
plt.scatter(X_pca2[y==lab, 0], X_pca2[y==lab, 1], label=iris.target_names[lab], color=col)
plt.xlabel('First principal component')
plt.ylabel('Second principal component')
plt.legend()
plt.title('Iris data after PCA')
plt.show()
# Let us check the actual numbers behind the principal components.
print('PCA components (directions):')
print(pca2.components_)
# Reconstruct original data from two components (inverse transform).
X_reconstructed = pca2.inverse_transform(X_pca2)
print('First few reconstructed rows:')
print(X_reconstructed[:5])
# Sometimes, PCA is useful before running other models.
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LogisticRegression
# Split the PCA data into train and test sets.
X_train, X_test, y_train, y_test = train_test_split(X_pca2, y, test_size=0.3, random_state=42)
# Fit a simple classifier on top of PCA components.
model = LogisticRegression()
model.fit(X_train, y_train)
score = model.score(X_test, y_test)
print('Test set accuracy:', score)
# Challenge: What if we use only 1 principal component?
pca1 = PCA(n_components=1)
X_pca1 = pca1.fit_transform(X_scaled)
print('First few values with one component:')
print(X_pca1[:5])
# Handling non-numeric data: PCA needs numbers only.
# Let us pretend one column is text to see what happens.
X_test_bad = X.copy()
X_test_bad['fake_text'] = ['flower']*len(X_test_bad)
try:
_ = scaler.fit_transform(X_test_bad)
except Exception as e:
print('Error:', e)
Mini-Project: Reducing a Real Dataset#
Now you try! Pick another dataset, or use Iris, and walk through these steps:
- Load the data.
- Standardize the features.
- Run PCA for 2-3 components.
- Visualize in a scatter plot.
Tip: Try the Wine dataset with sklearn or a CSV from the list above.
Recap#
- PCA helps us explore and simplify big sets of numbers.
- Always scale your features before running PCA.
- You can visualize main directions for data patterns.
- PCA is used in visualization, cleaner data, and even preprocessing for machine learning.
Try out PCA with your own data to see patterns appear!
Thank you for learning PCA with us! Please like and subscribe to our channel for more beginner Python and data science tutorials. Let us keep growing our coding skills together!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



