Lesson 9 · Python For Time Series
Understanding Probability Distributions in Python with SciPy for Statistical Analysis
In this lesson, you will use Python to explore random events and probability distributions using the SciPy library. You will see real data, make plots, and…
- CoursePython For Time Series
- Lesson9 of 30
- Video10 min
- FormatJupyter notebook · 17 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynb
Welcome to Python Probability Distributions with SciPy!#
In this lesson, you will use Python to explore random events and probability distributions using the SciPy library.
You will see real data, make plots, and run simple simulations.
Let us get started!
# Get ready! Import core libraries.
import warnings; warnings.filterwarnings("ignore")
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
from scipy import stats
What are Probability Distributions?#
Probability distributions let us model random events, like rolling dice or measuring heights.
With distributions, you can describe the chance of different outcomes.
SciPy makes it easy to use many probability distributions in Python.
# Simulate random dice rolls using a probability distribution
dice_rolls = np.random.choice([1, 2, 3, 4, 5, 6], size=10)
print("Ten random dice rolls:", dice_rolls)
# Data setup: real-world time series for later analysis
url = "https://raw.githubusercontent.com/jbrownlee/Datasets/master/shampoo.csv"
df = pd.read_csv(url)
print("Data shape:", df.shape)
print("First few rows:")
print(df.head())
# Plot the shampoo sales data to explore it visually
plt.figure(figsize=(8, 4))
plt.plot(df["Sales"], marker="o")
plt.title("Monthly Shampoo Sales (Time Series)")
plt.xlabel("Month")
plt.ylabel("Sales")
plt.show()
What is a Normal Distribution?#
A normal distribution is a classic 'bell curve' shape.
It describes many things you measure in the world, like heights or test scores.
Let us draw one in Python.
# Plotting a normal distribution (bell curve)
x = np.linspace(-4, 4, 100)
pdf = stats.norm.pdf(x, loc=0, scale=1)
plt.plot(x, pdf, label="Standard Normal PDF")
plt.title("Normal Distribution (Bell Curve)")
plt.xlabel("Value")
plt.ylabel("Probability Density")
plt.legend()
plt.show()
# Generate random heights with a normal distribution
np.random.seed(42)
heights = np.random.normal(loc=170, scale=10, size=1000)
plt.hist(heights, bins=30, color="skyblue", edgecolor="black")
plt.title("Simulated Heights (cm)")
plt.xlabel("Height (cm)")
plt.ylabel("Count")
plt.show()
# Find probability of a value under the normal curve
mean = 170
std = 10
prob = stats.norm.cdf(180, loc=mean, scale=std)
print(f"Chance a person is shorter than 180 cm: {prob:.2f}")
The Binomial Distribution#
Binomial distributions model how often something happens out of a number of tries, like flipping a coin ten times.
Let us see how likely it is to get certain results with ten coin flips.
# Compute and plot binomial probabilities
n = 10 # Number of flips
p = 0.5 # Chance of heads
x = np.arange(0, n+1)
pmf = stats.binom.pmf(x, n, p)
plt.bar(x, pmf, color="lightcoral", edgecolor="black")
plt.title("Coin Flips: Probability of Getting Heads")
plt.xlabel("Number of Heads")
plt.ylabel("Probability")
plt.show()
# What is the probability of getting exactly five heads?
prob_5 = stats.binom.pmf(5, n, p)
print(f"Probability of exactly 5 heads: {prob_5:.2f}")
# Practice: Simulate a person's answer to a question using input()
answer = input("What is your favorite random number between 1 and 6? ")
print("Thank you! Your random number is:", answer)
Practice Prompt#
Can you think of a real-life example where knowing probabilities could help make a decision?
Share your ideas or write them in your notebook!
# Use the Poisson distribution: rare events in sales (mini-project setup)
monthly_average = df["Sales"].mean()
poisson_sample = stats.poisson.rvs(monthly_average, size=12)
plt.bar(np.arange(1, 13), poisson_sample, color="goldenrod")
plt.title("Simulated Monthly Sales Events (Poisson)")
plt.xlabel("Month")
plt.ylabel("Simulated Sales Counts")
plt.show()
# Mini-project: Bootstrap confidence intervals for the mean (advanced!)
np.random.seed(42)
sample_means = [np.mean(np.random.choice(df["Sales"], size=20, replace=True)) for _ in range(1000)]
lower, upper = np.percentile(sample_means, [2.5, 97.5])
print(f"95% bootstrap confidence interval: ({lower:.2f}, {upper:.2f})")
# Best practices: Always check your data before modeling
print(df.isnull().sum())
print(df.dtypes)
# Troubleshooting: What if your values are not numbers?
try:
numbers = pd.to_numeric(df["Sales"])
except Exception as e:
print("Could not convert to numbers:", e)
# Extra tip: Compare two distributions side by side
group_a = np.random.normal(170, 10, 200)
group_b = np.random.normal(160, 8, 200)
plt.hist(group_a, bins=25, alpha=0.6, label="Group A")
plt.hist(group_b, bins=25, alpha=0.6, label="Group B")
plt.legend()
plt.title("Comparing Two Groups (Heights)")
plt.xlabel("Height (cm)")
plt.ylabel("Count")
plt.show()
Challenge Exercise#
Change the code above to see how the shape of the distribution changes for different values of average (mean) and spread (standard deviation).
Try different situations and see what happens!
Lesson Recap#
In this notebook, you learned to:
- Use probability distributions in Python with SciPy
- Simulate and plot random events
- Analyze real world data with probability models
- Practice coding with real datasets and exercises
Keep experimenting and trying new questions!
Thanks for Learning!#
If you enjoyed this lesson, please like and subscribe on YouTube.
Stay curious and keep exploring Python!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



