Mathew K Analytics

Lesson 10 · Python For Time Series

Understanding Correlation and Covariance in Python for Statistical Analysis

Welcome! In this lesson, we will explore two key concepts for analyzing how variables move together: correlation and covariance. You will learn the basics,…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb
 

Correlation and Covariance in Python#

Welcome! In this lesson, we will explore two key concepts for analyzing how variables move together: correlation and covariance.

You will learn the basics, see real-world data examples, and get tips for using these tools in your own projects.

Let's dive in!

What does correlation mean?#

Correlation shows how closely related two sets of numbers are.

If two things go up or down together, the correlation is strong. If one goes up and the other goes down, it is negative.

Correlation values range from -1 (always move opposite) to +1 (always move together). 0 means there is no clear link.

import warnings
warnings.filterwarnings("ignore")

# We will often use pandas and matplotlib for data tasks
import pandas as pd
import matplotlib.pyplot as plt
# Data setup: Let's load a real dataset about airline passengers over time
url = "https://raw.githubusercontent.com/jbrownlee/Datasets/master/airline-passengers.csv"
data = pd.read_csv(url)
print("Data shape:", data.shape)
data.head()
Data shape: (144, 2)
Month Passengers
0 1949-01 112
1 1949-02 118
2 1949-03 132
3 1949-04 129
4 1949-05 121
# Plot the passenger numbers over time
plt.figure(figsize=(10,4))
plt.plot(data['Month'], data['Passengers'], label="Passengers")
plt.xlabel("Month")
plt.ylabel("Number of Passengers")
plt.title("Monthly Airline Passengers")
plt.xticks(rotation=45)
plt.legend()
plt.tight_layout()
plt.show()
No description has been provided for this image

What is covariance?#

Covariance is a number that tells us if two sets of values move together and by how much.

Unlike correlation, covariance can be any size, including big positive or negative values.

Covariance tells us direction, but not scale.

# Let us build two small lists and see how covariance works
score_math = [55, 85, 67, 90, 77]
score_science = [60, 80, 70, 95, 79]

scores = pd.DataFrame({"Math": score_math, "Science": score_science})
print(scores)

print("Covariance matrix:")
print(scores.cov())
   Math  Science
0    55       60
1    85       80
2    67       70
3    90       95
4    77       79
Covariance matrix:
           Math  Science
Math     198.20   174.95
Science  174.95   168.70
# Correlation works very much like covariance but gives numbers between -1 and +1.
print("Correlation matrix:")
print(scores.corr())
Correlation matrix:
             Math   Science
Math     1.000000  0.956763
Science  0.956763  1.000000

Why do we standardize data for correlation?#

Covariance changes if you use bigger numbers, but correlation always fits between -1 and +1.

This makes correlation easier to compare between different problems.

# Try changing the scale. Multiply math scores by 10.
scores["Math_x10"] = scores["Math"] * 10

print(scores[["Math_x10", "Science"]].cov())
print(scores[["Math_x10", "Science"]].corr())
          Math_x10  Science
Math_x10   19820.0   1749.5
Science     1749.5    168.7
          Math_x10   Science
Math_x10  1.000000  0.956763
Science   0.956763  1.000000
# Correlation in a real-world dataset (airline passengers)
corr = data['Passengers'].corr(pd.Series(range(len(data))))
print(f"Correlation between 'Passengers' and months: {corr:.2f}")
Correlation between 'Passengers' and months: 0.92
# Compare two years: are 1958 and 1959 related?
data["Year"] = data["Month"].str[:4].astype(int)
yearly = data.groupby("Year")["Passengers"].sum()
print(yearly)
print("Covariance and correlation:")
print(yearly.cov(yearly.shift(1)))
print(yearly.corr(yearly.shift(1)))
Year
1949    1520
1950    1676
1951    2042
1952    2364
1953    2700
1954    2867
1955    3408
1956    3939
1957    4421
1958    4572
1959    5140
1960    5714
Name: Passengers, dtype: int64
Covariance and correlation:
1626100.0181818185
0.9937811388620401
# Missing data can sometimes happen in the real world
data_missing = data.copy()
data_missing.loc[0, "Passengers"] = None
data_missing.loc[1, "Passengers"] = None
print("Does pandas handle missing numbers?")
print(data_missing["Passengers"].corr(pd.Series(range(len(data_missing)))))
Does pandas handle missing numbers?
0.9220443835323521
# How to deal with missing data: fill or drop?
# Option 1: Fill missing with the average value
filled = data_missing["Passengers"].fillna(data_missing["Passengers"].mean())
print(filled.corr(pd.Series(range(len(filled)))))

# Option 2: Remove all missing rows
dropped = data_missing.dropna()
print(dropped["Passengers"].corr(pd.Series(range(len(dropped)))))
0.9029013619330383
0.9228440043127354
# Using input() to enter two lists and compare their correlation
user1 = input("List one, comma separated numbers: ")
user2 = input("List two, comma separated numbers: ")
lst1 = [float(x.strip()) for x in user1.split(",")]
lst2 = [float(x.strip()) for x in user2.split(",")]
print("Covariance:", pd.Series(lst1).cov(pd.Series(lst2)))
print("Correlation:", pd.Series(lst1).corr(pd.Series(lst2)))
Covariance: 5.0
Correlation: 0.9999999999999999
 
# Quick: Find the column pair in the airline data with the strongest correlation
numeric_data = data.select_dtypes(include='number')
print(numeric_data.corr())
            Passengers      Year
Passengers    1.000000  0.921824
Year          0.921824  1.000000
# Creating a scatter plot to see the relationship
plt.scatter(scores["Math"], scores["Science"], color='red')
plt.xlabel("Math scores")
plt.ylabel("Science scores")
plt.title("Scatter Plot: Math vs. Science")
plt.show()
No description has been provided for this image
# Easy tip: Get correlation fast with numpy
import numpy as np
np.corrcoef(score_math, score_science)
array([[1.        , 0.95676346],
       [0.95676346, 1.        ]])
# Mini project part 1: Shampoo sales vs. time
url2 = "https://raw.githubusercontent.com/jbrownlee/Datasets/master/shampoo.csv"
shampoo = pd.read_csv(url2)
shampoo.columns = ["Month", "Sales"]
shampoo["Sales"] = pd.to_numeric(shampoo["Sales"], errors="coerce")

print(shampoo.head())
plt.plot(shampoo['Sales'])
plt.title("Shampoo Sales Over Time")
plt.xlabel("Month")
plt.ylabel("Sales")
plt.show()
  Month  Sales
0  1-01  266.0
1  1-02  145.9
2  1-03  183.1
3  1-04  119.3
4  1-05  180.3
No description has been provided for this image
# Mini project part 2: Correlation between airline and shampoo trends
cor_air_shamp = pd.Series(data["Passengers"][:len(shampoo)]).corr(shampoo["Sales"])
print(f"Correlation between airline passenger and shampoo sales (first {len(shampoo)} rows): {cor_air_shamp:.2f}")
Correlation between airline passenger and shampoo sales (first 36 rows): 0.64

Common mistakes and troubleshooting#

  • Double-check your data for missing or text values instead of numbers.
  • Make sure your data lists or columns line up and have the same length.
  • Use dropna() or fillna() to handle any missing values before getting correlation.

Extra tips#

  • For big datasets, try .corr() on the whole table for quick exploration.
  • Plots like scatter or line charts help you spot outliers and patterns quickly.
  • For deep analysis, combine correlation with time shifting or differencing.
# Challenge: Can you find a pair with negative correlation?
lst_a = [10, 8, 6, 4, 2]
lst_b = [1, 3, 5, 7, 9]
print("Correlation:", pd.Series(lst_a).corr(pd.Series(lst_b)))
Correlation: -0.9999999999999999

Recap: What did we learn?#

  • Correlation and covariance show how two things move together.
  • Correlation is scale-free, covariance is not.
  • Real-world data often needs cleaning before analysis.
  • Plots help make relationships clear.
  • Try your own data next!

Thanks for learning!#

If you enjoyed this lesson, like and subscribe to our YouTube for more hands-on Python.

Keep exploring and let us know what you build next!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.