Mathew K Analytics

Lesson 7 · Statistics for data analysts

Hypothesis Testing Fundamentals | Statistics #7

Video seven of the 15-part series: what a p-value actually is, what it isn't, and the two ways a hypothesis test can go wrong. Real Sample Superstore order…

What you'll learn

Datasets used in this lesson

Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.

📓 Full notebook

Download .ipynb

Statistics for Data Analysts, Video 7: Hypothesis Testing Fundamentals#

  • Video seven of the 15-part series: what a p-value actually is, what it isn't, and the two ways a hypothesis test can go wrong.
  • Real Sample Superstore order data, plus a permutation test built from scratch.
  • Let's get into it.

Before You Start#

  • Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
  • You'll need pandas, NumPy, and SciPy.
  • Place superstore_sales.csv in the same folder as this notebook.
import pandas as pd
import numpy as np
from scipy import stats

df = pd.read_csv('superstore_sales.csv')
pop_mean = df['Sales'].mean()
print(f'Real population mean Sales: {pop_mean:.2f}')
Real population mean Sales: 229.86

Part 1: Setting Up a Hypothesis Test#

  • Null hypothesis (H0): The real average Sales for Central region orders equals the real overall population average.
  • Alternative hypothesis (H1): The real average Sales for Central region orders is different from the real overall population average.
  • This is a two-sided test: H1 allows for either higher or lower, not a specific direction.
central_sample = df[df['Region'] == 'Central'].sample(n=80, random_state=11)['Sales']
t_stat, p_value = stats.ttest_1samp(central_sample, pop_mean)
print(f'Real Central sample mean: {central_sample.mean():.2f}')
print(f'Real hypothesized mean (H0): {pop_mean:.2f}')
print(f'Real t-statistic: {t_stat:.3f}')
print(f'Real p-value: {p_value:.4f}')
Real Central sample mean: 135.32
Real hypothesized mean (H0): 229.86
Real t-statistic: -2.972
Real p-value: 0.0039

The real p-value here is the probability of seeing a sample mean at least this far from the hypothesized value, if H0 were actually true. It is small, well under the conventional 0.05 threshold, or significance level, alpha. That's evidence against H0, so we reject it: this real sample gives us reason to believe Central really does differ from the overall average, which lines up with what we already confirmed directly in video six using the full real population.

Part 2: What a P-Value Actually Is, Built from Scratch#

office_sample = df[df['Category'] == 'Office Supplies'].sample(n=100, random_state=3)['Sales'].values
tech_sample = df[df['Category'] == 'Technology'].sample(n=100, random_state=3)['Sales'].values
observed_diff = tech_sample.mean() - office_sample.mean()
print(f'Real observed difference (Technology - Office Supplies): {observed_diff:.2f}')
Real observed difference (Technology - Office Supplies): 487.29
rng = np.random.default_rng(seed=3)
combined = np.concatenate([office_sample, tech_sample])
n1 = len(office_sample)
perm_diffs = np.empty(10000)
for i in range(10000):
    shuffled = rng.permutation(combined)
    perm_diffs[i] = shuffled[n1:].mean() - shuffled[:n1].mean()
p_perm = (np.abs(perm_diffs) >= abs(observed_diff)).mean()
print(f'Real permutation p-value: {p_perm:.4f}')
Real permutation p-value: 0.0000
import matplotlib.pyplot as plt

plt.figure(figsize=(9, 4))
plt.hist(perm_diffs, bins=50, color='steelblue', edgecolor='white')
plt.axvline(observed_diff, color='darkred', linewidth=2, label=f'Real observed diff = {observed_diff:.1f}')
plt.title('Simulated Null Distribution (Permutation Test) vs. the Real Observed Difference')
plt.xlabel('Difference in Means Under Shuffled Labels')
plt.legend()
plt.show()
No description has been provided for this image
t_param, p_param = stats.ttest_ind(tech_sample, office_sample)
print(f'Real parametric t-test p-value: {p_param:.4f}')
print(f'Real permutation p-value: {p_perm:.4f}')
Real parametric t-test p-value: 0.0376
Real permutation p-value: 0.0000

Part 3: Type I and Type II Errors#

rng2 = np.random.default_rng(seed=42)
central_vals = df[df['Region'] == 'Central']['Sales'].values
false_positives = 0
trials = 2000
for _ in range(trials):
    a = rng2.choice(central_vals, size=60, replace=False)
    b = rng2.choice(central_vals, size=60, replace=False)
    _, p = stats.ttest_ind(a, b)
    if p < 0.05:
        false_positives += 1
print(f'Real Type I error rate (H0 genuinely true here): {false_positives / trials:.1%}')
Real Type I error rate (H0 genuinely true here): 3.8%
office_vals = df[df['Category'] == 'Office Supplies']['Sales'].values
tech_vals = df[df['Category'] == 'Technology']['Sales'].values
for n in [10, 30, 80, 200]:
    significant = 0
    trials2 = 1000
    for _ in range(trials2):
        a = rng2.choice(office_vals, size=n, replace=False)
        b = rng2.choice(tech_vals, size=n, replace=False)
        _, p = stats.ttest_ind(a, b)
        if p < 0.05:
            significant += 1
    power = significant / trials2
    print(f'n={n}: real statistical power = {power:.1%} (Type II error rate = {1 - power:.1%})')
n=10: real statistical power = 28.1% (Type II error rate = 71.9%)
n=30: real statistical power = 61.1% (Type II error rate = 38.9%)
n=80: real statistical power = 89.8% (Type II error rate = 10.2%)
n=200: real statistical power = 99.9% (Type II error rate = 0.1%)

Contrast this with West versus Central Sales from video six: that real difference is only about $11, tiny next to a standard deviation in the hundreds. Rerunning this exact power simulation on that real pair (try it yourself) shows power barely rising above alpha even at the full real sample size, because the true real effect is nearly swallowed by noise. Not every real, genuine difference is practically detectable, and knowing that distinction before collecting data is exactly what video eight's power calculations are for.

Part 4: One-Sided vs. Two-Sided Tests#

t_two, p_two = stats.ttest_ind(tech_sample, office_sample, alternative='two-sided')
t_one, p_one = stats.ttest_ind(tech_sample, office_sample, alternative='greater')
print(f'Two-sided H1 (Technology != Office Supplies): p = {p_two:.5f}')
print(f"One-sided H1 (Technology > Office Supplies): p = {p_one:.5f}")
Two-sided H1 (Technology != Office Supplies): p = 0.03759
One-sided H1 (Technology > Office Supplies): p = 0.01879

Part 5: Common Misinterpretations, Set Straight#

  • A p-value is not the probability that H0 is true. It's the probability of seeing data this extreme, assuming H0 already is true.
  • A small p-value is not proof of a large or important effect. The West-versus-Central case here has a real, statistically detectable-in-principle difference that's practically tiny; statistical significance and practical significance are separate questions.
  • Failing to reject H0 is not proof H0 is true. It might just mean the real sample was too small to detect a genuine effect, exactly what the Type II simulation in part three demonstrated directly.
  • Running many tests and reporting only the significant ones inflates the real false-positive rate far above the stated alpha; video ten covers exactly how to correct for that.

Wrap-Up: What You Learned#

  • Setting up null and alternative hypotheses, and running a one-sample t-test on real Central region data.
  • Building a p-value from scratch with a permutation test, and seeing exactly what 'assuming H0 is true' means as a simulated distribution.
  • Type I errors (false alarms) and Type II errors (missed real effects), both measured directly through simulation, including statistical power climbing with sample size.
  • One-sided versus two-sided tests, and why the choice must come before looking at the real data.
  • Four common p-value misinterpretations, grounded in real examples from this video and the last.
  • Video eight puts all of this into a full A/B test: sample size and power calculations planned in advance, run on clearly labeled synthetic experiment data since no real randomized trial exists in this dataset. Subscribe so it lands automatically see you there.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.