Lesson 8 · Statistics for data analysts
A/B Testing & Experiment Design That Works | Statistics #8
Video eight of the 15-part series: planning a test properly before running it, not just analyzing results after the fact. Important note up front: none of…
- CourseStatistics for data analysts
- Lesson8 of 15
- Video16 min
- FormatJupyter notebook · 9 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbStatistics for Data Analysts, Video 8: A/B Testing and Experiment Design#
- Video eight of the 15-part series: planning a test properly before running it, not just analyzing results after the fact.
- Important note up front: none of our sourced datasets contain a real randomized experiment log, so this video uses clearly labeled synthetic visitor data, generated with a fixed seed and named as synthetic throughout. Every other video in this series uses real data; this is the one deliberate exception, called out explicitly.
- Let's get into it.
Before You Start#
- Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
- Install statsmodels if you don't have it yet:
pip install statsmodels. - No dataset file is needed this video; every value is generated in the notebook itself.
import numpy as np
from scipy import stats
from statsmodels.stats.power import NormalIndPower
from statsmodels.stats.proportion import proportion_effectsize, proportions_ztest
print('Libraries loaded')
Part 1: Framing the Experiment#
- Null hypothesis (H0): The new checkout button's true conversion rate equals the current button's true conversion rate.
- Alternative hypothesis (H1): The two true conversion rates differ.
- Baseline (control) conversion rate: assumed 10%, a reasonable synthetic starting point for this scenario.
- Minimum detectable effect: the team wants to reliably detect a rise to 12%, a 2 percentage point, or 20% relative, lift.
Part 2: Sample Size and Power, Calculated in Advance#
baseline_p = 0.10
target_p = 0.12
effect_size = proportion_effectsize(target_p, baseline_p)
print(f'Cohen\'s h effect size: {effect_size:.4f}')
power_analysis = NormalIndPower()
required_n = power_analysis.solve_power(effect_size=effect_size, alpha=0.05, power=0.8, ratio=1, alternative='two-sided')
required_n = int(np.ceil(required_n))
print(f'Required sample size per group: {required_n}')
print(f'Total visitors needed: {required_n * 2}')
for target in [0.11, 0.12, 0.15]:
eff = proportion_effectsize(target, baseline_p)
n_req = int(np.ceil(power_analysis.solve_power(effect_size=eff, alpha=0.05, power=0.8, ratio=1, alternative='two-sided')))
print(f'To detect a rise to {target:.0%}: need {n_req} per group')
Part 3: Generating the Synthetic Experiment and Analyzing It#
rng = np.random.default_rng(seed=2024)
n_per_group = required_n
true_p_control = 0.10
true_p_treatment = 0.12
synthetic_control = rng.binomial(1, true_p_control, size=n_per_group)
synthetic_treatment = rng.binomial(1, true_p_treatment, size=n_per_group)
print(f'Synthetic control conversions: {synthetic_control.sum()} of {n_per_group} ({synthetic_control.mean():.4f})')
print(f'Synthetic treatment conversions: {synthetic_treatment.sum()} of {n_per_group} ({synthetic_treatment.mean():.4f})')
count = np.array([synthetic_treatment.sum(), synthetic_control.sum()])
nobs = np.array([n_per_group, n_per_group])
z_stat, p_value = proportions_ztest(count, nobs, alternative='two-sided')
print(f'z-statistic: {z_stat:.4f}')
print(f'p-value: {p_value:.4f}')
rate_control = synthetic_control.mean()
rate_treatment = synthetic_treatment.mean()
diff = rate_treatment - rate_control
se_diff = np.sqrt(rate_control * (1 - rate_control) / n_per_group + rate_treatment * (1 - rate_treatment) / n_per_group)
z_crit = stats.norm.ppf(0.975)
ci_low, ci_high = diff - z_crit * se_diff, diff + z_crit * se_diff
print(f'Observed lift: {diff:.4f} ({diff / rate_control:.1%} relative)')
print(f'95% CI for the lift: ({ci_low:.4f}, {ci_high:.4f})')
This particular synthetic run did not come back statistically significant, even though the true underlying treatment rate really was higher, by design. That's not a bug in the analysis, it's exactly what an 80%-power test is supposed to do about 20% of the time: correctly-designed experiments still miss a real effect on any single run. Part four proves this by rerunning the exact same synthetic setup a thousand times.
Part 4: Confirming the Power Calculation Was Right#
rng2 = np.random.default_rng(seed=2024)
sig_count = 0
trials = 1000
for _ in range(trials):
c = rng2.binomial(1, true_p_control, size=n_per_group)
t = rng2.binomial(1, true_p_treatment, size=n_per_group)
cnt = np.array([t.sum(), c.sum()])
nb = np.array([n_per_group, n_per_group])
_, p = proportions_ztest(cnt, nb, alternative='two-sided')
if p < 0.05:
sig_count += 1
print(f'Synthetic empirical power: {sig_count / trials:.1%} (planned target was 80%)')
Part 5: The Peeking Problem#
rng3 = np.random.default_rng(seed=55)
true_p_no_effect = 0.10
check_points = np.arange(200, n_per_group + 1, 200)
trials2 = 1000
peeking_false_positives = 0
fixed_false_positives = 0
for _ in range(trials2):
c_full = rng3.binomial(1, true_p_no_effect, size=n_per_group)
t_full = rng3.binomial(1, true_p_no_effect, size=n_per_group)
peeked_sig = False
for cp in check_points:
cnt = np.array([t_full[:cp].sum(), c_full[:cp].sum()])
nb = np.array([cp, cp])
_, p_peek = proportions_ztest(cnt, nb, alternative='two-sided')
if p_peek < 0.05:
peeked_sig = True
break
if peeked_sig:
peeking_false_positives += 1
cnt_final = np.array([t_full.sum(), c_full.sum()])
nb_final = np.array([n_per_group, n_per_group])
_, p_fixed = proportions_ztest(cnt_final, nb_final, alternative='two-sided')
if p_fixed < 0.05:
fixed_false_positives += 1
print(f'False positive rate WITH daily peeking: {peeking_false_positives / trials2:.1%}')
print(f'False positive rate checking ONLY at the planned end: {fixed_false_positives / trials2:.1%}')
Peeking pushes the true false-positive rate well above the intended 5%, even though every single peek individually used a correct 5%-alpha test. The problem is cumulative: each additional look is another chance for random noise to cross the significance line, and stopping as soon as it does cherry-picks exactly those lucky moments. The fix is what video seven's Type I error section already implied: decide the sample size in advance, as done in part two, and look at the result exactly once.
Wrap-Up: What You Learned#
- Framing an A/B test as a formal hypothesis test, with an explicit baseline rate and minimum detectable effect.
- Calculating required sample size and power in advance with statsmodels, before collecting a single data point.
- Analyzing a synthetic two-proportion experiment with a z-test and a confidence interval for the lift, including an honest non-significant result.
- Confirming that non-significant result was expected, by empirically reproducing the planned power rate across 1,000 synthetic replications.
- The peeking problem: repeated early looks at accumulating data genuinely inflate the false-positive rate far above the intended alpha.
- A reminder that this video's data was clearly labeled synthetic throughout; video nine returns to real data with non-parametric tests for when the normal-based assumptions behind a t-test or z-test don't hold. Subscribe so it lands automatically see you there.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



