Mathew K Analytics

Lesson 22 · Probability and Statistics in python

Understanding Covariance and Correlation: Pearson and Spearman Explained

Welcome! Today we will learn how variables relate to each other. We will explore covariance and two ways to measure correlation. You do not need advanced…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Lesson: Covariance and Correlation (Pearson and Spearman)#

Welcome! Today we will learn how variables relate to each other. We will explore covariance and two ways to measure correlation.

You do not need advanced math to get started.

Let us begin!

What are Covariance and Correlation?#

  • Covariance shows if two variables move together or in opposite directions.
  • Correlation tells how strong and in what direction that relationship is.
  • We measure relationships using numbers from data.

Let us see some simple examples!

import warnings
from statsmodels.tools.sm_exceptions import ConvergenceWarning, ValueWarning
warnings.filterwarnings("ignore", category=ValueWarning)
warnings.filterwarnings("ignore", category=ConvergenceWarning)
warnings.filterwarnings("ignore", category=RuntimeWarning)
warnings.filterwarnings('ignore')  # Suppress all warnings

import numpy as np
import pandas as pd
import seaborn as sns
import matplotlib.pyplot as plt  # For visualizations

# Data setup: Let us use the MPG dataset
df = sns.load_dataset('mpg')
print('MPG dataset shape:', df.shape)
df = df.dropna(subset=['mpg', 'horsepower', 'weight'])  # Remove missing values for this lesson
print('Preview:')
df = df.reset_index(drop=True)
df.head()
MPG dataset shape: (398, 9)
Preview:
mpg cylinders displacement horsepower weight acceleration model_year origin name
0 18.0 8 307.0 130.0 3504 12.0 70 usa chevrolet chevelle malibu
1 15.0 8 350.0 165.0 3693 11.5 70 usa buick skylark 320
2 18.0 8 318.0 150.0 3436 11.0 70 usa plymouth satellite
3 16.0 8 304.0 150.0 3433 12.0 70 usa amc rebel sst
4 17.0 8 302.0 140.0 3449 10.5 70 usa ford torino

Visual Exploration#

A scatterplot is a great way to see relationships between variables.

Let us plot 'horsepower' versus 'mpg' (miles per gallon).

What do you notice?

plt.figure(figsize=(6,4))
sns.scatterplot(x='horsepower', y='mpg', data=df)
plt.title('Horsepower vs MPG')
plt.xlabel('Horsepower')
plt.ylabel('Miles per Gallon (mpg)')
plt.show()
No description has been provided for this image

Intuitive Idea#

If two variables move together, the points go in a line.

  • High horsepower, low mpg? That is a negative relationship.
  • High horsepower, high mpg? That would be positive.

Now let us see this with numbers.

# Calculate covariance between horsepower and mpg
cov_hp_mpg = np.cov(df['horsepower'], df['mpg'], ddof=0)[0,1]
print(f"Covariance between horsepower and mpg: {cov_hp_mpg:.2f}")
Covariance between horsepower and mpg: -233.26

What does the number mean?#

  • Covariance is hard to compare because it depends on the units of data.
  • Large values can mean little if scales are different.
  • That is why we use correlation.

Correlation puts everything on the same scale.

# Calculate Pearson correlation between horsepower and mpg
pearson_corr = df['horsepower'].corr(df['mpg'])
print(f"Pearson correlation between horsepower and mpg: {pearson_corr:.2f}")
Pearson correlation between horsepower and mpg: -0.78
# Calculate Spearman correlation between horsepower and mpg
spearman_corr = df['horsepower'].corr(df['mpg'], method='spearman')
print(f"Spearman correlation between horsepower and mpg: {spearman_corr:.2f}")
Spearman correlation between horsepower and mpg: -0.85

Quick Reference Table#

Name Measures Shape Required Use Case
Covariance Movement (raw) Any First look, math
Pearson Linear relation Line-shaped Most numeric relationships
Spearman Rank (monotonic) Any curve Rankings, non-linear trends

Tip: Monotonic means always up or always down, not zig-zaggy.

# Heatmap of Pearson correlations for all numeric columns
corr_matrix = df.corr(numeric_only=True)
plt.figure(figsize=(8,6))
sns.heatmap(corr_matrix, annot=True, cmap='coolwarm', fmt='.2f')
plt.title('Pearson Correlation Heatmap')
plt.show()
No description has been provided for this image
# What if variables are ranks, not numbers? Let us try Spearman on car origin vs mpg.
df['origin_cat'] = df['origin'].astype('category').cat.codes
origin_mpg_spearman = df['origin_cat'].corr(df['mpg'], method='spearman')
print(f"Spearman correlation between origin (as code) and mpg: {origin_mpg_spearman:.2f}")
Spearman correlation between origin (as code) and mpg: -0.53
# Visualize the relationship with a regression line using seaborn
sns.lmplot(x='horsepower', y='mpg', data=df, aspect=1.2, scatter_kws={'alpha':0.5})
plt.title('Horsepower vs MPG with Trend Line')
plt.xlabel('Horsepower')
plt.ylabel('Miles per Gallon (mpg)')
plt.show()
No description has been provided for this image
# Outliers can strongly affect correlation. Let us check what happens if we add a wild point.
df2 = df.copy()
columns = list(df2.columns)
# Make a row that fits the number/order of columns. Most columns are set to None except what we want to set.
new_row = [None] * len(columns)
if 'mpg' in columns: new_row[columns.index('mpg')] = 5
if 'cylinders' in columns: new_row[columns.index('cylinders')] = 8
if 'horsepower' in columns: new_row[columns.index('horsepower')] = 500
if 'weight' in columns: new_row[columns.index('weight')] = 8000
if 'origin' in columns:
    try:
        cat_origin = df2['origin'].cat.categories.get_loc('usa') if hasattr(df2['origin'], 'cat') else 0
    except Exception:
        cat_origin = 0
    new_row[columns.index('origin')] = 'usa'
else:
    cat_origin = 0
if 'origin_cat' in columns:
    new_row[columns.index('origin_cat')] = cat_origin

# Add the new row
df2.loc[len(df2)] = new_row

pearson_outlier = df2['horsepower'].corr(df2['mpg'])
print(f"Pearson correlation after adding outlier: {pearson_outlier:.2f}")
Pearson correlation after adding outlier: -0.74

Summary: When to Use Each#

  • Use covariance to describe same-direction movement (but numbers are hard to compare).
  • Use Pearson for line-style relationships without dramatic outliers.
  • Use Spearman for rankings, or if not sure the relationship is straight.

If there are wild values or categories, pick Spearman.

Next, let us practice!

# Practice: Try your own columns!
col1 = input("Enter name of first numeric column (e.g., horsepower): ")
col2 = input("Enter name of second numeric column (e.g., mpg): ")
try:
    if col1 in df.columns and col2 in df.columns:
        # Confirm columns are numeric
        c1 = pd.to_numeric(df[col1], errors='coerce')
        c2 = pd.to_numeric(df[col2], errors='coerce')
        mask = ~(c1.isna() | c2.isna())
        print('Covariance:', np.cov(c1[mask], c2[mask], ddof=0)[0,1])
        print('Pearson:', c1[mask].corr(c2[mask]))
        print('Spearman:', c1[mask].corr(c2[mask], method='spearman'))
    else:
        print("One or both column names not found.")
except Exception as e:
    print('Error:', e)
    
Covariance: -5503.36560027072
Pearson: -0.8322442148315754
Spearman: -0.8755851198739869

Troubleshooting Tips#

  • Blank or missing values give errors. Use dropna to fix before running correlation.
  • Use the exact column names: Try df.columns to check them.
  • If you get a value near zero: variables might not be related.
  • Large outliers can break Pearson. Try Spearman instead.

You are getting better at finding relationships!

# Extra tip: Quickly see which columns relate to mpg using Pearson
try:
    mpg_corrs = df.select_dtypes(include=[np.number]).corr()['mpg'].drop('mpg').sort_values()
    print('Correlation with mpg:')
    print(mpg_corrs)
except Exception as e:
    print(f'Error: {e}')
    
Correlation with mpg:
weight         -0.832244
displacement   -0.805127
horsepower     -0.778427
cylinders      -0.777618
origin_cat     -0.474800
acceleration    0.423329
model_year      0.580541
Name: mpg, dtype: float64
# Challenge: Try a subset of the data, such as only cars from Japan
japan_cars = df[df['origin'] == 'japan']
if len(japan_cars) > 0:
    print('Pearson correlation (horsepower vs mpg, Japan only):', japan_cars['horsepower'].corr(japan_cars['mpg']))
    print('Spearman:', japan_cars['horsepower'].corr(japan_cars['mpg'], method='spearman'))
else:
    print('No data for Japan!')
    
Pearson correlation (horsepower vs mpg, Japan only): -0.6730950429373183
Spearman: -0.7008569239486375

Recap#

You have learned:

  • Covariance shows 'do they move together?'
  • Pearson: best for lines, sensitive to big outliers.
  • Spearman: best for ranks and non-linear trends, handles wild values better.

Try this lesson with your own data!

Thanks for learning with us.

Curious for more?#

Post your own findings or questions in the comments below! Subscribe to the channel for more beginner-friendly statistics lessons.

See you in the next video!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.