Lesson 14 · Statistics for data analysts
Bayesian Statistics Basics for Analysts | Statistics #14
Video fourteen of the 15-part series: a different lens on the same problems, updating belief with evidence instead of just testing against a null. Real…
- CourseStatistics for data analysts
- Lesson14 of 15
- Video15 min
- FormatJupyter notebook · 9 code cells
- Data1 dataset
What you'll learn
Datasets used in this lesson
Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.
- mushroom_predictions.csv78.3 KB
📓 Full notebook
Download .ipynbStatistics for Data Analysts, Video 14: Bayesian Statistics Basics#
- Video fourteen of the 15-part series: a different lens on the same problems, updating belief with evidence instead of just testing against a null.
- Real mushroom data for Bayes' theorem itself, then a return to the clearly labeled synthetic A/B test data from video eight for the Bayesian-versus-frequentist comparison.
- Let's get into it.
Before You Start#
- Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
- You'll need pandas, NumPy, SciPy, and Matplotlib.
- Place mushroom_predictions.csv in the same folder as this notebook.
import pandas as pd
import numpy as np
from scipy import stats
import matplotlib.pyplot as plt
mushrooms = pd.read_csv('mushroom_predictions.csv')
print('Libraries and real mushroom data loaded')
Part 1: Bayes' Theorem, Verified on Real Data#
mushrooms['odor_5'] = (mushrooms['odor'] == 5).astype(int)
mushrooms['edible'] = (mushrooms['actual'] == 'e').astype(int)
p_edible = mushrooms['edible'].mean()
p_odor5 = mushrooms['odor_5'].mean()
p_odor5_given_edible = mushrooms.loc[mushrooms['edible'] == 1, 'odor_5'].mean()
print(f'Real P(edible): {p_edible:.4f}')
print(f'Real P(odor code 5): {p_odor5:.4f}')
print(f'Real P(odor code 5 | edible): {p_odor5_given_edible:.4f}')
bayes_result = p_odor5_given_edible * p_edible / p_odor5
direct_result = mushrooms.loc[mushrooms['odor_5'] == 1, 'edible'].mean()
print(f'Real P(edible | odor code 5) via Bayes\' theorem: {bayes_result:.4f}')
print(f'Real P(edible | odor code 5) computed directly: {direct_result:.4f}')
Part 2: Priors, Likelihood, and Posteriors#
rng = np.random.default_rng(seed=2024)
n_per_group = 3835
true_p_control = 0.10
true_p_treatment = 0.12
synthetic_control = rng.binomial(1, true_p_control, size=n_per_group)
synthetic_treatment = rng.binomial(1, true_p_treatment, size=n_per_group)
conv_control = synthetic_control.sum()
conv_treatment = synthetic_treatment.sum()
print(f'Synthetic control: {conv_control} of {n_per_group} conversions')
print(f'Synthetic treatment: {conv_treatment} of {n_per_group} conversions')
prior_a, prior_b = 1, 1
post_a_control = prior_a + conv_control
post_b_control = prior_b + (n_per_group - conv_control)
post_a_treatment = prior_a + conv_treatment
post_b_treatment = prior_b + (n_per_group - conv_treatment)
print(f'Control posterior: Beta({post_a_control}, {post_b_control}), mean={post_a_control / (post_a_control + post_b_control):.4f}')
print(f'Treatment posterior: Beta({post_a_treatment}, {post_b_treatment}), mean={post_a_treatment / (post_a_treatment + post_b_treatment):.4f}')
Part 3: Watching the Posterior Update#
checkpoints = [0, 100, 500, 1500, n_per_group]
x_vals = np.linspace(0, 0.25, 500)
plt.figure(figsize=(9, 5))
for cp in checkpoints:
a = prior_a + synthetic_treatment[:cp].sum()
b = prior_b + (cp - synthetic_treatment[:cp].sum())
plt.plot(x_vals, stats.beta.pdf(x_vals, a, b), label=f'n={cp}')
plt.axvline(true_p_treatment, color='black', linestyle='--', label='True rate (0.12)')
plt.title('Posterior Distribution for the Treatment Rate as Synthetic Data Accumulates')
plt.xlabel('Conversion Rate')
plt.ylabel('Density')
plt.legend()
plt.show()
Part 4: Credible Intervals vs. Confidence Intervals#
credible_low, credible_high = stats.beta.ppf([0.025, 0.975], post_a_treatment, post_b_treatment)
print(f'Real 95% credible interval for the treatment rate: ({credible_low:.4f}, {credible_high:.4f})')
rate_treatment = conv_treatment / n_per_group
se_treatment = np.sqrt(rate_treatment * (1 - rate_treatment) / n_per_group)
z_crit = stats.norm.ppf(0.975)
ci_low, ci_high = rate_treatment - z_crit * se_treatment, rate_treatment + z_crit * se_treatment
print(f'Real 95% confidence interval for the treatment rate: ({ci_low:.4f}, {ci_high:.4f})')
The numbers nearly match, but the meaning doesn't. The credible interval genuinely says: given this real data and this prior, there's a 95% probability the true treatment rate lies in this range, a direct statement about the parameter. The confidence interval, as video six covered in detail, says something about the long-run behavior of the procedure across repeated real sampling, not a direct probability statement about this one interval. Numerically similar here; philosophically, not the same claim.
Part 5: A Bayesian Answer to the A/B Test#
rng2 = np.random.default_rng(seed=99)
samples_control = rng2.beta(post_a_control, post_b_control, size=200000)
samples_treatment = rng2.beta(post_a_treatment, post_b_treatment, size=200000)
p_treatment_better = (samples_treatment > samples_control).mean()
print(f'Real P(treatment rate > control rate | data): {p_treatment_better:.1%}')
lift_samples = samples_treatment - samples_control
credible_lift_low, credible_lift_high = np.percentile(lift_samples, [2.5, 97.5])
print(f'Real posterior mean lift: {lift_samples.mean():.4f}')
print(f'Real 95% credible interval for the lift: ({credible_lift_low:.4f}, {credible_lift_high:.4f})')
Wrap-Up: What You Learned#
- Bayes' theorem, verified two independent ways on real mushroom data.
- Priors, likelihood, and posteriors, using a Beta-Binomial conjugate model on the same clearly labeled synthetic A/B test data from video eight.
- Watching a posterior distribution sharpen and shift as synthetic evidence accumulates.
- Credible intervals versus confidence intervals: numerically similar under an uninformative prior, but different in what they actually claim.
- A Bayesian probability-of-improvement answer for the same A/B test that came back non-significant under a frequentist p-value, illustrating how the two frameworks can offer complementary, not contradictory, information.
- Video fifteen closes the series with a full capstone: a real, complete statistical analysis combining distribution checks, confidence intervals, hypothesis tests, and regression into one written report. Subscribe so it lands automatically see you there.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



