Mathew K Analytics

Lesson 8 · Statistics for data analysts

A/B Testing & Experiment Design That Works | Statistics #8

Video eight of the 15-part series: planning a test properly before running it, not just analyzing results after the fact. Important note up front: none of…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Statistics for Data Analysts, Video 8: A/B Testing and Experiment Design#

  • Video eight of the 15-part series: planning a test properly before running it, not just analyzing results after the fact.
  • Important note up front: none of our sourced datasets contain a real randomized experiment log, so this video uses clearly labeled synthetic visitor data, generated with a fixed seed and named as synthetic throughout. Every other video in this series uses real data; this is the one deliberate exception, called out explicitly.
  • Let's get into it.

Before You Start#

  • Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
  • Install statsmodels if you don't have it yet: pip install statsmodels.
  • No dataset file is needed this video; every value is generated in the notebook itself.
import numpy as np
from scipy import stats
from statsmodels.stats.power import NormalIndPower
from statsmodels.stats.proportion import proportion_effectsize, proportions_ztest
print('Libraries loaded')
Libraries loaded

Part 1: Framing the Experiment#

  • Null hypothesis (H0): The new checkout button's true conversion rate equals the current button's true conversion rate.
  • Alternative hypothesis (H1): The two true conversion rates differ.
  • Baseline (control) conversion rate: assumed 10%, a reasonable synthetic starting point for this scenario.
  • Minimum detectable effect: the team wants to reliably detect a rise to 12%, a 2 percentage point, or 20% relative, lift.

Part 2: Sample Size and Power, Calculated in Advance#

baseline_p = 0.10
target_p = 0.12
effect_size = proportion_effectsize(target_p, baseline_p)
print(f'Cohen\'s h effect size: {effect_size:.4f}')
Cohen's h effect size: 0.0640
power_analysis = NormalIndPower()
required_n = power_analysis.solve_power(effect_size=effect_size, alpha=0.05, power=0.8, ratio=1, alternative='two-sided')
required_n = int(np.ceil(required_n))
print(f'Required sample size per group: {required_n}')
print(f'Total visitors needed: {required_n * 2}')
Required sample size per group: 3835
Total visitors needed: 7670
for target in [0.11, 0.12, 0.15]:
    eff = proportion_effectsize(target, baseline_p)
    n_req = int(np.ceil(power_analysis.solve_power(effect_size=eff, alpha=0.05, power=0.8, ratio=1, alternative='two-sided')))
    print(f'To detect a rise to {target:.0%}: need {n_req} per group')
To detect a rise to 11%: need 14745 per group
To detect a rise to 12%: need 3835 per group
To detect a rise to 15%: need 681 per group

Part 3: Generating the Synthetic Experiment and Analyzing It#

rng = np.random.default_rng(seed=2024)
n_per_group = required_n
true_p_control = 0.10
true_p_treatment = 0.12
synthetic_control = rng.binomial(1, true_p_control, size=n_per_group)
synthetic_treatment = rng.binomial(1, true_p_treatment, size=n_per_group)
print(f'Synthetic control conversions: {synthetic_control.sum()} of {n_per_group} ({synthetic_control.mean():.4f})')
print(f'Synthetic treatment conversions: {synthetic_treatment.sum()} of {n_per_group} ({synthetic_treatment.mean():.4f})')
Synthetic control conversions: 397 of 3835 (0.1035)
Synthetic treatment conversions: 425 of 3835 (0.1108)
count = np.array([synthetic_treatment.sum(), synthetic_control.sum()])
nobs = np.array([n_per_group, n_per_group])
z_stat, p_value = proportions_ztest(count, nobs, alternative='two-sided')
print(f'z-statistic: {z_stat:.4f}')
print(f'p-value: {p_value:.4f}')
z-statistic: 1.0336
p-value: 0.3013
rate_control = synthetic_control.mean()
rate_treatment = synthetic_treatment.mean()
diff = rate_treatment - rate_control
se_diff = np.sqrt(rate_control * (1 - rate_control) / n_per_group + rate_treatment * (1 - rate_treatment) / n_per_group)
z_crit = stats.norm.ppf(0.975)
ci_low, ci_high = diff - z_crit * se_diff, diff + z_crit * se_diff
print(f'Observed lift: {diff:.4f} ({diff / rate_control:.1%} relative)')
print(f'95% CI for the lift: ({ci_low:.4f}, {ci_high:.4f})')
Observed lift: 0.0073 (7.1% relative)
95% CI for the lift: (-0.0065, 0.0211)

This particular synthetic run did not come back statistically significant, even though the true underlying treatment rate really was higher, by design. That's not a bug in the analysis, it's exactly what an 80%-power test is supposed to do about 20% of the time: correctly-designed experiments still miss a real effect on any single run. Part four proves this by rerunning the exact same synthetic setup a thousand times.

Part 4: Confirming the Power Calculation Was Right#

rng2 = np.random.default_rng(seed=2024)
sig_count = 0
trials = 1000
for _ in range(trials):
    c = rng2.binomial(1, true_p_control, size=n_per_group)
    t = rng2.binomial(1, true_p_treatment, size=n_per_group)
    cnt = np.array([t.sum(), c.sum()])
    nb = np.array([n_per_group, n_per_group])
    _, p = proportions_ztest(cnt, nb, alternative='two-sided')
    if p < 0.05:
        sig_count += 1
print(f'Synthetic empirical power: {sig_count / trials:.1%} (planned target was 80%)')
Synthetic empirical power: 80.1% (planned target was 80%)

Part 5: The Peeking Problem#

rng3 = np.random.default_rng(seed=55)
true_p_no_effect = 0.10
check_points = np.arange(200, n_per_group + 1, 200)
trials2 = 1000
peeking_false_positives = 0
fixed_false_positives = 0
for _ in range(trials2):
    c_full = rng3.binomial(1, true_p_no_effect, size=n_per_group)
    t_full = rng3.binomial(1, true_p_no_effect, size=n_per_group)
    peeked_sig = False
    for cp in check_points:
        cnt = np.array([t_full[:cp].sum(), c_full[:cp].sum()])
        nb = np.array([cp, cp])
        _, p_peek = proportions_ztest(cnt, nb, alternative='two-sided')
        if p_peek < 0.05:
            peeked_sig = True
            break
    if peeked_sig:
        peeking_false_positives += 1
    cnt_final = np.array([t_full.sum(), c_full.sum()])
    nb_final = np.array([n_per_group, n_per_group])
    _, p_fixed = proportions_ztest(cnt_final, nb_final, alternative='two-sided')
    if p_fixed < 0.05:
        fixed_false_positives += 1
print(f'False positive rate WITH daily peeking: {peeking_false_positives / trials2:.1%}')
print(f'False positive rate checking ONLY at the planned end: {fixed_false_positives / trials2:.1%}')
False positive rate WITH daily peeking: 23.2%
False positive rate checking ONLY at the planned end: 4.8%

Peeking pushes the true false-positive rate well above the intended 5%, even though every single peek individually used a correct 5%-alpha test. The problem is cumulative: each additional look is another chance for random noise to cross the significance line, and stopping as soon as it does cherry-picks exactly those lucky moments. The fix is what video seven's Type I error section already implied: decide the sample size in advance, as done in part two, and look at the result exactly once.

Wrap-Up: What You Learned#

  • Framing an A/B test as a formal hypothesis test, with an explicit baseline rate and minimum detectable effect.
  • Calculating required sample size and power in advance with statsmodels, before collecting a single data point.
  • Analyzing a synthetic two-proportion experiment with a z-test and a confidence interval for the lift, including an honest non-significant result.
  • Confirming that non-significant result was expected, by empirically reproducing the planned power rate across 1,000 synthetic replications.
  • The peeking problem: repeated early looks at accumulating data genuinely inflate the false-positive rate far above the intended alpha.
  • A reminder that this video's data was clearly labeled synthetic throughout; video nine returns to real data with non-parametric tests for when the normal-based assumptions behind a t-test or z-test don't hold. Subscribe so it lands automatically see you there.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.