Lesson 2 · Statistics for data analysts
Discrete Distributions Explained: Binomial & Poisson | Statistics #2
Video two of the 15-part series: formalizing the probability rules from video one into two workhorse discrete distributions, the binomial and the Poisson.…
- CourseStatistics for data analysts
- Lesson2 of 15
- Video16 min
- FormatJupyter notebook · 12 code cells
- Data1 dataset
What you'll learn
Datasets used in this lesson
Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.
- superstore_sales.csv715.5 KB
📓 Full notebook
Download .ipynbStatistics for Data Analysts, Video 2: Discrete Distributions#
- Video two of the 15-part series: formalizing the probability rules from video one into two workhorse discrete distributions, the binomial and the Poisson.
- Real Sample Superstore order data throughout: real Technology-order rates and real daily order counts.
- Let's get into it.
Before You Start#
- Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
- Install SciPy if you don't have it yet:
pip install scipy. - Place superstore_sales.csv in the same folder as this notebook.
import pandas as pd
import numpy as np
from scipy import stats
import matplotlib.pyplot as plt
df = pd.read_csv('superstore_sales.csv')
df['Order Date'] = pd.to_datetime(df['Order Date'])
print(df.shape)
Part 1: The Binomial Distribution#
p_tech = (df['Category'] == 'Technology').mean()
n = 20
print(f'Real P(Technology) per order line: {p_tech:.4f}')
print(f'Modeling batches of n={n} order lines')
rng = np.random.default_rng(seed=7)
n_batches = 10000
simulated_counts = rng.binomial(n=n, p=p_tech, size=n_batches)
print(f'Simulated mean Technology count per batch: {simulated_counts.mean():.3f}')
print(f'Theoretical mean (n*p): {n * p_tech:.3f}')
k_values = np.arange(0, n + 1)
pmf = stats.binom.pmf(k_values, n, p_tech)
for k, prob in zip(k_values[:6], pmf[:6]):
print(f'P(exactly {k} Technology orders in a batch of {n}) = {prob:.4f}')
plt.figure(figsize=(9, 4))
plt.bar(k_values, pmf, color='steelblue', alpha=0.7, label='Theoretical PMF')
sim_probs = pd.Series(simulated_counts).value_counts(normalize=True).sort_index()
plt.scatter(sim_probs.index, sim_probs.values, color='darkred', zorder=3, label='Simulated frequencies')
plt.title(f'Binomial(n={n}, p={p_tech:.3f}): Theoretical vs. Simulated')
plt.xlabel('Number of Technology Orders in Batch')
plt.ylabel('Probability')
plt.legend()
plt.show()
Part 2: Binomial Mean, Variance, and the CDF#
mean_binom, var_binom = stats.binom.stats(n, p_tech, moments='mv')
print(f'Theoretical mean: {float(mean_binom):.3f}')
print(f'Theoretical variance: {float(var_binom):.3f}')
print(f'Simulated mean: {simulated_counts.mean():.3f}, simulated variance: {simulated_counts.var():.3f}')
p_at_most_2 = stats.binom.cdf(2, n, p_tech)
p_at_least_5 = 1 - stats.binom.cdf(4, n, p_tech)
print(f'P(at most 2 Technology orders) = {p_at_most_2:.4f}')
print(f'P(at least 5 Technology orders) = {p_at_least_5:.4f}')
Part 3: The Poisson Distribution#
daily_counts = df.groupby(df['Order Date'].dt.date).size()
lam = daily_counts.mean()
print(f'Number of real days with at least one order: {len(daily_counts)}')
print(f'Real average daily order-line count (lambda): {lam:.3f}')
k_values2 = np.arange(0, daily_counts.max() + 1)
poisson_pmf = stats.poisson.pmf(k_values2, lam)
observed_freq = daily_counts.value_counts(normalize=True).sort_index()
plt.figure(figsize=(9, 4))
plt.bar(k_values2, poisson_pmf, color='steelblue', alpha=0.7, label=f'Poisson(lambda={lam:.2f}) PMF')
plt.scatter(observed_freq.index, observed_freq.values, color='darkred', zorder=3, label='Real observed frequencies')
plt.title('Real Daily Order Counts vs. a Fitted Poisson')
plt.xlabel('Order Lines per Day')
plt.ylabel('Probability')
plt.legend()
plt.show()
var_real = daily_counts.var()
print(f'Real daily count mean: {lam:.3f}')
print(f'Real daily count variance: {var_real:.3f}')
print(f'Poisson assumes mean == variance; real ratio (variance/mean) = {var_real / lam:.2f}')
This kind of overdispersion, real variance exceeding the real mean, is extremely common in business event data and is usually a sign the event rate itself varies (by weekday, by season, by promotion) rather than staying constant like the Poisson assumes. The usual next step is a negative binomial model, which adds a second parameter to absorb that extra variance; that's outside today's scope, but worth knowing the name for when a plain Poisson underfits real data like this.
Part 4: Poisson Probabilities in Practice#
p_zero_or_one = stats.poisson.cdf(1, lam)
p_busy_day = 1 - stats.poisson.cdf(15, lam)
print(f'Modeled P(0 or 1 order lines in a day) = {p_zero_or_one:.4f}')
print(f'Modeled P(more than 15 order lines in a day) = {p_busy_day:.4f}')
real_p_busy = (daily_counts > 15).mean()
print(f'Real observed P(more than 15 order lines) = {real_p_busy:.4f}')
Part 5: Choosing Between Binomial and Poisson#
- Binomial: a fixed, known number of trials (n), each independently succeeding or failing with the same probability (p). Use it when you can name n up front, like a batch of 20 order lines.
- Poisson: events happening at some average rate over a continuous interval, with no natural upper bound on the count. Use it when there's no fixed n, like however many order lines happen to land on a given day.
- As n grows large and p gets small while np stays roughly fixed, the binomial distribution actually converges to a Poisson distribution with lambda = np, a useful approximation in practice.
big_n, small_p = 2000, p_tech / 100
approx_lambda = big_n * small_p
binom_pmf_tail = stats.binom.pmf(np.arange(0, 8), big_n, small_p)
poisson_pmf_tail = stats.poisson.pmf(np.arange(0, 8), approx_lambda)
comparison = pd.DataFrame({'k': np.arange(0, 8), 'Binomial': binom_pmf_tail, 'Poisson_approx': poisson_pmf_tail})
print(comparison.round(5))
Wrap-Up: What You Learned#
- The binomial distribution, for counting successes across a fixed number of independent trials, verified against real simulated batches.
- The Poisson distribution, for counting events over an interval, fit to real daily order counts, including an honest look at overdispersion.
- Mean, variance, PMF, and CDF for both distributions, computed with scipy.stats.
- When to reach for each distribution, and how the binomial approximates the Poisson under the right conditions.
- Video three moves to continuous distributions: the normal, uniform, and exponential, fit to real mtcars and stock-return data. Subscribe so it lands automatically see you there.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



