Lesson 22 · Probability and Statistics in python
Understanding Covariance and Correlation: Pearson and Spearman Explained
Welcome! Today we will learn how variables relate to each other. We will explore covariance and two ways to measure correlation. You do not need advanced…
- CourseProbability and Statistics in python
- Lesson22 of 35
- Video13 min
- FormatJupyter notebook · 12 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbLesson: Covariance and Correlation (Pearson and Spearman)#
Welcome! Today we will learn how variables relate to each other. We will explore covariance and two ways to measure correlation.
You do not need advanced math to get started.
Let us begin!
What are Covariance and Correlation?#
- Covariance shows if two variables move together or in opposite directions.
- Correlation tells how strong and in what direction that relationship is.
- We measure relationships using numbers from data.
Let us see some simple examples!
import warnings
from statsmodels.tools.sm_exceptions import ConvergenceWarning, ValueWarning
warnings.filterwarnings("ignore", category=ValueWarning)
warnings.filterwarnings("ignore", category=ConvergenceWarning)
warnings.filterwarnings("ignore", category=RuntimeWarning)
warnings.filterwarnings('ignore') # Suppress all warnings
import numpy as np
import pandas as pd
import seaborn as sns
import matplotlib.pyplot as plt # For visualizations
# Data setup: Let us use the MPG dataset
df = sns.load_dataset('mpg')
print('MPG dataset shape:', df.shape)
df = df.dropna(subset=['mpg', 'horsepower', 'weight']) # Remove missing values for this lesson
print('Preview:')
df = df.reset_index(drop=True)
df.head()
Visual Exploration#
A scatterplot is a great way to see relationships between variables.
Let us plot 'horsepower' versus 'mpg' (miles per gallon).
What do you notice?
plt.figure(figsize=(6,4))
sns.scatterplot(x='horsepower', y='mpg', data=df)
plt.title('Horsepower vs MPG')
plt.xlabel('Horsepower')
plt.ylabel('Miles per Gallon (mpg)')
plt.show()
Intuitive Idea#
If two variables move together, the points go in a line.
- High horsepower, low mpg? That is a negative relationship.
- High horsepower, high mpg? That would be positive.
Now let us see this with numbers.
# Calculate covariance between horsepower and mpg
cov_hp_mpg = np.cov(df['horsepower'], df['mpg'], ddof=0)[0,1]
print(f"Covariance between horsepower and mpg: {cov_hp_mpg:.2f}")
What does the number mean?#
- Covariance is hard to compare because it depends on the units of data.
- Large values can mean little if scales are different.
- That is why we use correlation.
Correlation puts everything on the same scale.
# Calculate Pearson correlation between horsepower and mpg
pearson_corr = df['horsepower'].corr(df['mpg'])
print(f"Pearson correlation between horsepower and mpg: {pearson_corr:.2f}")
# Calculate Spearman correlation between horsepower and mpg
spearman_corr = df['horsepower'].corr(df['mpg'], method='spearman')
print(f"Spearman correlation between horsepower and mpg: {spearman_corr:.2f}")
Quick Reference Table#
| Name | Measures | Shape Required | Use Case |
|---|---|---|---|
| Covariance | Movement (raw) | Any | First look, math |
| Pearson | Linear relation | Line-shaped | Most numeric relationships |
| Spearman | Rank (monotonic) | Any curve | Rankings, non-linear trends |
Tip: Monotonic means always up or always down, not zig-zaggy.
# Heatmap of Pearson correlations for all numeric columns
corr_matrix = df.corr(numeric_only=True)
plt.figure(figsize=(8,6))
sns.heatmap(corr_matrix, annot=True, cmap='coolwarm', fmt='.2f')
plt.title('Pearson Correlation Heatmap')
plt.show()
# What if variables are ranks, not numbers? Let us try Spearman on car origin vs mpg.
df['origin_cat'] = df['origin'].astype('category').cat.codes
origin_mpg_spearman = df['origin_cat'].corr(df['mpg'], method='spearman')
print(f"Spearman correlation between origin (as code) and mpg: {origin_mpg_spearman:.2f}")
# Visualize the relationship with a regression line using seaborn
sns.lmplot(x='horsepower', y='mpg', data=df, aspect=1.2, scatter_kws={'alpha':0.5})
plt.title('Horsepower vs MPG with Trend Line')
plt.xlabel('Horsepower')
plt.ylabel('Miles per Gallon (mpg)')
plt.show()
# Outliers can strongly affect correlation. Let us check what happens if we add a wild point.
df2 = df.copy()
columns = list(df2.columns)
# Make a row that fits the number/order of columns. Most columns are set to None except what we want to set.
new_row = [None] * len(columns)
if 'mpg' in columns: new_row[columns.index('mpg')] = 5
if 'cylinders' in columns: new_row[columns.index('cylinders')] = 8
if 'horsepower' in columns: new_row[columns.index('horsepower')] = 500
if 'weight' in columns: new_row[columns.index('weight')] = 8000
if 'origin' in columns:
try:
cat_origin = df2['origin'].cat.categories.get_loc('usa') if hasattr(df2['origin'], 'cat') else 0
except Exception:
cat_origin = 0
new_row[columns.index('origin')] = 'usa'
else:
cat_origin = 0
if 'origin_cat' in columns:
new_row[columns.index('origin_cat')] = cat_origin
# Add the new row
df2.loc[len(df2)] = new_row
pearson_outlier = df2['horsepower'].corr(df2['mpg'])
print(f"Pearson correlation after adding outlier: {pearson_outlier:.2f}")
Summary: When to Use Each#
- Use covariance to describe same-direction movement (but numbers are hard to compare).
- Use Pearson for line-style relationships without dramatic outliers.
- Use Spearman for rankings, or if not sure the relationship is straight.
If there are wild values or categories, pick Spearman.
Next, let us practice!
# Practice: Try your own columns!
col1 = input("Enter name of first numeric column (e.g., horsepower): ")
col2 = input("Enter name of second numeric column (e.g., mpg): ")
try:
if col1 in df.columns and col2 in df.columns:
# Confirm columns are numeric
c1 = pd.to_numeric(df[col1], errors='coerce')
c2 = pd.to_numeric(df[col2], errors='coerce')
mask = ~(c1.isna() | c2.isna())
print('Covariance:', np.cov(c1[mask], c2[mask], ddof=0)[0,1])
print('Pearson:', c1[mask].corr(c2[mask]))
print('Spearman:', c1[mask].corr(c2[mask], method='spearman'))
else:
print("One or both column names not found.")
except Exception as e:
print('Error:', e)
Troubleshooting Tips#
- Blank or missing values give errors. Use dropna to fix before running correlation.
- Use the exact column names: Try df.columns to check them.
- If you get a value near zero: variables might not be related.
- Large outliers can break Pearson. Try Spearman instead.
You are getting better at finding relationships!
# Extra tip: Quickly see which columns relate to mpg using Pearson
try:
mpg_corrs = df.select_dtypes(include=[np.number]).corr()['mpg'].drop('mpg').sort_values()
print('Correlation with mpg:')
print(mpg_corrs)
except Exception as e:
print(f'Error: {e}')
# Challenge: Try a subset of the data, such as only cars from Japan
japan_cars = df[df['origin'] == 'japan']
if len(japan_cars) > 0:
print('Pearson correlation (horsepower vs mpg, Japan only):', japan_cars['horsepower'].corr(japan_cars['mpg']))
print('Spearman:', japan_cars['horsepower'].corr(japan_cars['mpg'], method='spearman'))
else:
print('No data for Japan!')
Recap#
You have learned:
- Covariance shows 'do they move together?'
- Pearson: best for lines, sensitive to big outliers.
- Spearman: best for ranks and non-linear trends, handles wild values better.
Try this lesson with your own data!
Thanks for learning with us.
Curious for more?#
Post your own findings or questions in the comments below! Subscribe to the channel for more beginner-friendly statistics lessons.
See you in the next video!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



