Lesson 10 · Python For Time Series
Understanding Correlation and Covariance in Python for Statistical Analysis
Welcome! In this lesson, we will explore two key concepts for analyzing how variables move together: correlation and covariance. You will learn the basics,…
- CoursePython For Time Series
- Lesson10 of 30
- Video14 min
- FormatJupyter notebook · 19 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynb
Correlation and Covariance in Python#
Welcome! In this lesson, we will explore two key concepts for analyzing how variables move together: correlation and covariance.
You will learn the basics, see real-world data examples, and get tips for using these tools in your own projects.
Let's dive in!
What does correlation mean?#
Correlation shows how closely related two sets of numbers are.
If two things go up or down together, the correlation is strong. If one goes up and the other goes down, it is negative.
Correlation values range from -1 (always move opposite) to +1 (always move together). 0 means there is no clear link.
import warnings
warnings.filterwarnings("ignore")
# We will often use pandas and matplotlib for data tasks
import pandas as pd
import matplotlib.pyplot as plt
# Data setup: Let's load a real dataset about airline passengers over time
url = "https://raw.githubusercontent.com/jbrownlee/Datasets/master/airline-passengers.csv"
data = pd.read_csv(url)
print("Data shape:", data.shape)
data.head()
# Plot the passenger numbers over time
plt.figure(figsize=(10,4))
plt.plot(data['Month'], data['Passengers'], label="Passengers")
plt.xlabel("Month")
plt.ylabel("Number of Passengers")
plt.title("Monthly Airline Passengers")
plt.xticks(rotation=45)
plt.legend()
plt.tight_layout()
plt.show()
What is covariance?#
Covariance is a number that tells us if two sets of values move together and by how much.
Unlike correlation, covariance can be any size, including big positive or negative values.
Covariance tells us direction, but not scale.
# Let us build two small lists and see how covariance works
score_math = [55, 85, 67, 90, 77]
score_science = [60, 80, 70, 95, 79]
scores = pd.DataFrame({"Math": score_math, "Science": score_science})
print(scores)
print("Covariance matrix:")
print(scores.cov())
# Correlation works very much like covariance but gives numbers between -1 and +1.
print("Correlation matrix:")
print(scores.corr())
Why do we standardize data for correlation?#
Covariance changes if you use bigger numbers, but correlation always fits between -1 and +1.
This makes correlation easier to compare between different problems.
# Try changing the scale. Multiply math scores by 10.
scores["Math_x10"] = scores["Math"] * 10
print(scores[["Math_x10", "Science"]].cov())
print(scores[["Math_x10", "Science"]].corr())
# Correlation in a real-world dataset (airline passengers)
corr = data['Passengers'].corr(pd.Series(range(len(data))))
print(f"Correlation between 'Passengers' and months: {corr:.2f}")
# Compare two years: are 1958 and 1959 related?
data["Year"] = data["Month"].str[:4].astype(int)
yearly = data.groupby("Year")["Passengers"].sum()
print(yearly)
print("Covariance and correlation:")
print(yearly.cov(yearly.shift(1)))
print(yearly.corr(yearly.shift(1)))
# Missing data can sometimes happen in the real world
data_missing = data.copy()
data_missing.loc[0, "Passengers"] = None
data_missing.loc[1, "Passengers"] = None
print("Does pandas handle missing numbers?")
print(data_missing["Passengers"].corr(pd.Series(range(len(data_missing)))))
# How to deal with missing data: fill or drop?
# Option 1: Fill missing with the average value
filled = data_missing["Passengers"].fillna(data_missing["Passengers"].mean())
print(filled.corr(pd.Series(range(len(filled)))))
# Option 2: Remove all missing rows
dropped = data_missing.dropna()
print(dropped["Passengers"].corr(pd.Series(range(len(dropped)))))
# Using input() to enter two lists and compare their correlation
user1 = input("List one, comma separated numbers: ")
user2 = input("List two, comma separated numbers: ")
lst1 = [float(x.strip()) for x in user1.split(",")]
lst2 = [float(x.strip()) for x in user2.split(",")]
print("Covariance:", pd.Series(lst1).cov(pd.Series(lst2)))
print("Correlation:", pd.Series(lst1).corr(pd.Series(lst2)))
# Quick: Find the column pair in the airline data with the strongest correlation
numeric_data = data.select_dtypes(include='number')
print(numeric_data.corr())
# Creating a scatter plot to see the relationship
plt.scatter(scores["Math"], scores["Science"], color='red')
plt.xlabel("Math scores")
plt.ylabel("Science scores")
plt.title("Scatter Plot: Math vs. Science")
plt.show()
# Easy tip: Get correlation fast with numpy
import numpy as np
np.corrcoef(score_math, score_science)
# Mini project part 1: Shampoo sales vs. time
url2 = "https://raw.githubusercontent.com/jbrownlee/Datasets/master/shampoo.csv"
shampoo = pd.read_csv(url2)
shampoo.columns = ["Month", "Sales"]
shampoo["Sales"] = pd.to_numeric(shampoo["Sales"], errors="coerce")
print(shampoo.head())
plt.plot(shampoo['Sales'])
plt.title("Shampoo Sales Over Time")
plt.xlabel("Month")
plt.ylabel("Sales")
plt.show()
# Mini project part 2: Correlation between airline and shampoo trends
cor_air_shamp = pd.Series(data["Passengers"][:len(shampoo)]).corr(shampoo["Sales"])
print(f"Correlation between airline passenger and shampoo sales (first {len(shampoo)} rows): {cor_air_shamp:.2f}")
Common mistakes and troubleshooting#
- Double-check your data for missing or text values instead of numbers.
- Make sure your data lists or columns line up and have the same length.
- Use dropna() or fillna() to handle any missing values before getting correlation.
Extra tips#
- For big datasets, try .corr() on the whole table for quick exploration.
- Plots like scatter or line charts help you spot outliers and patterns quickly.
- For deep analysis, combine correlation with time shifting or differencing.
# Challenge: Can you find a pair with negative correlation?
lst_a = [10, 8, 6, 4, 2]
lst_b = [1, 3, 5, 7, 9]
print("Correlation:", pd.Series(lst_a).corr(pd.Series(lst_b)))
Recap: What did we learn?#
- Correlation and covariance show how two things move together.
- Correlation is scale-free, covariance is not.
- Real-world data often needs cleaning before analysis.
- Plots help make relationships clear.
- Try your own data next!
Thanks for learning!#
If you enjoyed this lesson, like and subscribe to our YouTube for more hands-on Python.
Keep exploring and let us know what you build next!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



