Mathew K Analytics

Lesson 68 · Python Fundamentals

Understanding Principal Component Analysis (PCA) in Python: A Comprehensive Tutorial

Principal Component Analysis, or PCA, is a popular way to understand and simplify complex datasets. Let us discover how PCA helps us find patterns, reduce…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb
 

Welcome to Principal Component Analysis (PCA) in Python#

Principal Component Analysis, or PCA, is a popular way to understand and simplify complex datasets. Let us discover how PCA helps us find patterns, reduce dimensions, and visualize data. We will use real examples, including the famous Iris flower dataset.

What is PCA?#

PCA finds patterns in data by looking for the directions where the data varies the most. It lets us make a big, complicated table of numbers into a small and simple oneso humans and computers can see important shapes.

# Let us set up Python to hide unnecessary warnings.
import warnings
warnings.filterwarnings('ignore')
# Data setup: Let us load and preview the Iris dataset.
from sklearn.datasets import load_iris
import pandas as pd

iris = load_iris()
X = pd.DataFrame(iris.data, columns=iris.feature_names)
y = pd.Series(iris.target, name='species')
 
print('Shape:', X.shape)
print('First 5 rows:')
print(X.head())
Shape: (150, 4)
First 5 rows:
   sepal length (cm)  sepal width (cm)  petal length (cm)  petal width (cm)
0                5.1               3.5                1.4               0.2
1                4.9               3.0                1.4               0.2
2                4.7               3.2                1.3               0.2
3                4.6               3.1                1.5               0.2
4                5.0               3.6                1.4               0.2
# Let us check how many samples are in each species group.
print('Species counts:')
print(y.value_counts())
Species counts:
species
0    50
1    50
2    50
Name: count, dtype: int64
# Let us view basic statistics for the features.
print(X.describe())
       sepal length (cm)  sepal width (cm)  petal length (cm)  \
count         150.000000        150.000000         150.000000   
mean            5.843333          3.057333           3.758000   
std             0.828066          0.435866           1.765298   
min             4.300000          2.000000           1.000000   
25%             5.100000          2.800000           1.600000   
50%             5.800000          3.000000           4.350000   
75%             6.400000          3.300000           5.100000   
max             7.900000          4.400000           6.900000   

       petal width (cm)  
count        150.000000  
mean           1.199333  
std            0.762238  
min            0.100000  
25%            0.300000  
50%            1.300000  
75%            1.800000  
max            2.500000  
# Standardize the data to mean 0 and variance 1.
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)

print('Mean of each column after scaling:')
print(X_scaled.mean(axis=0))
Mean of each column after scaling:
[-1.69031455e-15 -1.84297022e-15 -1.69864123e-15 -1.40924309e-15]

Starting PCA#

Let us start finding the important directions in our data. We will use scikit-learn, a well-known Python library, to do the math for us.

# Fit a PCA model and look at explained variance.
from sklearn.decomposition import PCA
pca = PCA()
pca.fit(X_scaled)

print('Explained variance ratio:', pca.explained_variance_ratio_)
Explained variance ratio: [0.72962445 0.22850762 0.03668922 0.00517871]
# Let us plot the cumulative variance explained.
import matplotlib.pyplot as plt
plt.plot(range(1, 1 + len(pca.explained_variance_ratio_)), pca.explained_variance_ratio_.cumsum(), marker='o')
plt.xlabel('Number of components')
plt.ylabel('Cumulative explained variance')
plt.title('Variance explained by PCA components')
plt.show()
No description has been provided for this image
# Transform data using two principal components for visualization.
pca2 = PCA(n_components=2)
X_pca2 = pca2.fit_transform(X_scaled)

print('Shape after PCA:', X_pca2.shape)
print('First few PCA-transformed rows:')
print(X_pca2[:5])
Shape after PCA: (150, 2)
First few PCA-transformed rows:
[[-2.26470281  0.4800266 ]
 [-2.08096115 -0.67413356]
 [-2.36422905 -0.34190802]
 [-2.29938422 -0.59739451]
 [-2.38984217  0.64683538]]
# Visualize the PCA result with a scatter plot.
plt.figure(figsize=(6,4))
for lab, col in zip(range(3), ['red', 'green', 'blue']):
    plt.scatter(X_pca2[y==lab, 0], X_pca2[y==lab, 1], label=iris.target_names[lab], color=col)
plt.xlabel('First principal component')
plt.ylabel('Second principal component')
plt.legend()
plt.title('Iris data after PCA')
plt.show()
No description has been provided for this image
# Let us check the actual numbers behind the principal components.
print('PCA components (directions):')
print(pca2.components_)
PCA components (directions):
[[ 0.52106591 -0.26934744  0.5804131   0.56485654]
 [ 0.37741762  0.92329566  0.02449161  0.06694199]]
# Reconstruct original data from two components (inverse transform).
X_reconstructed = pca2.inverse_transform(X_pca2)
print('First few reconstructed rows:')
print(X_reconstructed[:5])
First few reconstructed rows:
[[-0.99888895  1.05319838 -1.30270654 -1.24709825]
 [-1.33874781 -0.06192302 -1.22432772 -1.22057235]
 [-1.36096129  0.32111685 -1.38060338 -1.35833824]
 [-1.42359795  0.0677615  -1.34922386 -1.33881298]
 [-1.00113823  1.24091818 -1.37125365 -1.30661752]]
# Sometimes, PCA is useful before running other models.
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LogisticRegression

# Split the PCA data into train and test sets.
X_train, X_test, y_train, y_test = train_test_split(X_pca2, y, test_size=0.3, random_state=42)

# Fit a simple classifier on top of PCA components.
model = LogisticRegression()
model.fit(X_train, y_train)
score = model.score(X_test, y_test)
print('Test set accuracy:', score)
Test set accuracy: 0.9111111111111111
# Challenge: What if we use only 1 principal component?
pca1 = PCA(n_components=1)
X_pca1 = pca1.fit_transform(X_scaled)
print('First few values with one component:')
print(X_pca1[:5])
First few values with one component:
[[-2.26470281]
 [-2.08096115]
 [-2.36422905]
 [-2.29938422]
 [-2.38984217]]
# Handling non-numeric data: PCA needs numbers only.
# Let us pretend one column is text to see what happens.
X_test_bad = X.copy()
X_test_bad['fake_text'] = ['flower']*len(X_test_bad)
try:
    _ = scaler.fit_transform(X_test_bad)
except Exception as e:
    print('Error:', e)
    
Error: could not convert string to float: 'flower'

Mini-Project: Reducing a Real Dataset#

Now you try! Pick another dataset, or use Iris, and walk through these steps:

  1. Load the data.
  2. Standardize the features.
  3. Run PCA for 2-3 components.
  4. Visualize in a scatter plot.

Tip: Try the Wine dataset with sklearn or a CSV from the list above.

Recap#

  • PCA helps us explore and simplify big sets of numbers.
  • Always scale your features before running PCA.
  • You can visualize main directions for data patterns.
  • PCA is used in visualization, cleaner data, and even preprocessing for machine learning.

Try out PCA with your own data to see patterns appear!

Thank you for learning PCA with us! Please like and subscribe to our channel for more beginner Python and data science tutorials. Let us keep growing our coding skills together!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.