Mathew K Analytics

Lesson 2 · Statistics for data analysts

Discrete Distributions Explained: Binomial & Poisson | Statistics #2

Video two of the 15-part series: formalizing the probability rules from video one into two workhorse discrete distributions, the binomial and the Poisson.…

What you'll learn

Datasets used in this lesson

Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.

📓 Full notebook

Download .ipynb

Statistics for Data Analysts, Video 2: Discrete Distributions#

  • Video two of the 15-part series: formalizing the probability rules from video one into two workhorse discrete distributions, the binomial and the Poisson.
  • Real Sample Superstore order data throughout: real Technology-order rates and real daily order counts.
  • Let's get into it.

Before You Start#

  • Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
  • Install SciPy if you don't have it yet: pip install scipy.
  • Place superstore_sales.csv in the same folder as this notebook.
import pandas as pd
import numpy as np
from scipy import stats
import matplotlib.pyplot as plt

df = pd.read_csv('superstore_sales.csv')
df['Order Date'] = pd.to_datetime(df['Order Date'])
print(df.shape)
(9994, 9)

Part 1: The Binomial Distribution#

p_tech = (df['Category'] == 'Technology').mean()
n = 20
print(f'Real P(Technology) per order line: {p_tech:.4f}')
print(f'Modeling batches of n={n} order lines')
Real P(Technology) per order line: 0.1848
Modeling batches of n=20 order lines
rng = np.random.default_rng(seed=7)
n_batches = 10000
simulated_counts = rng.binomial(n=n, p=p_tech, size=n_batches)
print(f'Simulated mean Technology count per batch: {simulated_counts.mean():.3f}')
print(f'Theoretical mean (n*p): {n * p_tech:.3f}')
Simulated mean Technology count per batch: 3.699
Theoretical mean (n*p): 3.696
k_values = np.arange(0, n + 1)
pmf = stats.binom.pmf(k_values, n, p_tech)
for k, prob in zip(k_values[:6], pmf[:6]):
    print(f'P(exactly {k} Technology orders in a batch of {n}) = {prob:.4f}')
P(exactly 0 Technology orders in a batch of 20) = 0.0168
P(exactly 1 Technology orders in a batch of 20) = 0.0761
P(exactly 2 Technology orders in a batch of 20) = 0.1640
P(exactly 3 Technology orders in a batch of 20) = 0.2231
P(exactly 4 Technology orders in a batch of 20) = 0.2150
P(exactly 5 Technology orders in a batch of 20) = 0.1559
plt.figure(figsize=(9, 4))
plt.bar(k_values, pmf, color='steelblue', alpha=0.7, label='Theoretical PMF')
sim_probs = pd.Series(simulated_counts).value_counts(normalize=True).sort_index()
plt.scatter(sim_probs.index, sim_probs.values, color='darkred', zorder=3, label='Simulated frequencies')
plt.title(f'Binomial(n={n}, p={p_tech:.3f}): Theoretical vs. Simulated')
plt.xlabel('Number of Technology Orders in Batch')
plt.ylabel('Probability')
plt.legend()
plt.show()
No description has been provided for this image

Part 2: Binomial Mean, Variance, and the CDF#

mean_binom, var_binom = stats.binom.stats(n, p_tech, moments='mv')
print(f'Theoretical mean: {float(mean_binom):.3f}')
print(f'Theoretical variance: {float(var_binom):.3f}')
print(f'Simulated mean: {simulated_counts.mean():.3f}, simulated variance: {simulated_counts.var():.3f}')
Theoretical mean: 3.696
Theoretical variance: 3.013
Simulated mean: 3.699, simulated variance: 3.038
p_at_most_2 = stats.binom.cdf(2, n, p_tech)
p_at_least_5 = 1 - stats.binom.cdf(4, n, p_tech)
print(f'P(at most 2 Technology orders) = {p_at_most_2:.4f}')
print(f'P(at least 5 Technology orders) = {p_at_least_5:.4f}')
P(at most 2 Technology orders) = 0.2569
P(at least 5 Technology orders) = 0.3050

Part 3: The Poisson Distribution#

daily_counts = df.groupby(df['Order Date'].dt.date).size()
lam = daily_counts.mean()
print(f'Number of real days with at least one order: {len(daily_counts)}')
print(f'Real average daily order-line count (lambda): {lam:.3f}')
Number of real days with at least one order: 1237
Real average daily order-line count (lambda): 8.079
k_values2 = np.arange(0, daily_counts.max() + 1)
poisson_pmf = stats.poisson.pmf(k_values2, lam)
observed_freq = daily_counts.value_counts(normalize=True).sort_index()
plt.figure(figsize=(9, 4))
plt.bar(k_values2, poisson_pmf, color='steelblue', alpha=0.7, label=f'Poisson(lambda={lam:.2f}) PMF')
plt.scatter(observed_freq.index, observed_freq.values, color='darkred', zorder=3, label='Real observed frequencies')
plt.title('Real Daily Order Counts vs. a Fitted Poisson')
plt.xlabel('Order Lines per Day')
plt.ylabel('Probability')
plt.legend()
plt.show()
No description has been provided for this image
var_real = daily_counts.var()
print(f'Real daily count mean: {lam:.3f}')
print(f'Real daily count variance: {var_real:.3f}')
print(f'Poisson assumes mean == variance; real ratio (variance/mean) = {var_real / lam:.2f}')
Real daily count mean: 8.079
Real daily count variance: 39.047
Poisson assumes mean == variance; real ratio (variance/mean) = 4.83

This kind of overdispersion, real variance exceeding the real mean, is extremely common in business event data and is usually a sign the event rate itself varies (by weekday, by season, by promotion) rather than staying constant like the Poisson assumes. The usual next step is a negative binomial model, which adds a second parameter to absorb that extra variance; that's outside today's scope, but worth knowing the name for when a plain Poisson underfits real data like this.

Part 4: Poisson Probabilities in Practice#

p_zero_or_one = stats.poisson.cdf(1, lam)
p_busy_day = 1 - stats.poisson.cdf(15, lam)
print(f'Modeled P(0 or 1 order lines in a day) = {p_zero_or_one:.4f}')
print(f'Modeled P(more than 15 order lines in a day) = {p_busy_day:.4f}')
real_p_busy = (daily_counts > 15).mean()
print(f'Real observed P(more than 15 order lines) = {real_p_busy:.4f}')
Modeled P(0 or 1 order lines in a day) = 0.0028
Modeled P(more than 15 order lines in a day) = 0.0090
Real observed P(more than 15 order lines) = 0.1302

Part 5: Choosing Between Binomial and Poisson#

  • Binomial: a fixed, known number of trials (n), each independently succeeding or failing with the same probability (p). Use it when you can name n up front, like a batch of 20 order lines.
  • Poisson: events happening at some average rate over a continuous interval, with no natural upper bound on the count. Use it when there's no fixed n, like however many order lines happen to land on a given day.
  • As n grows large and p gets small while np stays roughly fixed, the binomial distribution actually converges to a Poisson distribution with lambda = np, a useful approximation in practice.
big_n, small_p = 2000, p_tech / 100
approx_lambda = big_n * small_p
binom_pmf_tail = stats.binom.pmf(np.arange(0, 8), big_n, small_p)
poisson_pmf_tail = stats.poisson.pmf(np.arange(0, 8), approx_lambda)
comparison = pd.DataFrame({'k': np.arange(0, 8), 'Binomial': binom_pmf_tail, 'Poisson_approx': poisson_pmf_tail})
print(comparison.round(5))
   k  Binomial  Poisson_approx
0  0   0.02473         0.02482
1  1   0.09159         0.09173
2  2   0.16949         0.16953
3  3   0.20900         0.20887
4  4   0.19320         0.19301
5  5   0.14280         0.14268
6  6   0.08791         0.08790
7  7   0.04637         0.04641

Wrap-Up: What You Learned#

  • The binomial distribution, for counting successes across a fixed number of independent trials, verified against real simulated batches.
  • The Poisson distribution, for counting events over an interval, fit to real daily order counts, including an honest look at overdispersion.
  • Mean, variance, PMF, and CDF for both distributions, computed with scipy.stats.
  • When to reach for each distribution, and how the binomial approximates the Poisson under the right conditions.
  • Video three moves to continuous distributions: the normal, uniform, and exponential, fit to real mtcars and stock-return data. Subscribe so it lands automatically see you there.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.