Lesson 7 · Statistics for data analysts
Hypothesis Testing Fundamentals | Statistics #7
Video seven of the 15-part series: what a p-value actually is, what it isn't, and the two ways a hypothesis test can go wrong. Real Sample Superstore order…
- CourseStatistics for data analysts
- Lesson7 of 15
- Video16 min
- FormatJupyter notebook · 9 code cells
- Data1 dataset
What you'll learn
Datasets used in this lesson
Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.
- superstore_sales.csv715.5 KB
📓 Full notebook
Download .ipynbStatistics for Data Analysts, Video 7: Hypothesis Testing Fundamentals#
- Video seven of the 15-part series: what a p-value actually is, what it isn't, and the two ways a hypothesis test can go wrong.
- Real Sample Superstore order data, plus a permutation test built from scratch.
- Let's get into it.
Before You Start#
- Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
- You'll need pandas, NumPy, and SciPy.
- Place superstore_sales.csv in the same folder as this notebook.
import pandas as pd
import numpy as np
from scipy import stats
df = pd.read_csv('superstore_sales.csv')
pop_mean = df['Sales'].mean()
print(f'Real population mean Sales: {pop_mean:.2f}')
Part 1: Setting Up a Hypothesis Test#
- Null hypothesis (H0): The real average Sales for Central region orders equals the real overall population average.
- Alternative hypothesis (H1): The real average Sales for Central region orders is different from the real overall population average.
- This is a two-sided test: H1 allows for either higher or lower, not a specific direction.
central_sample = df[df['Region'] == 'Central'].sample(n=80, random_state=11)['Sales']
t_stat, p_value = stats.ttest_1samp(central_sample, pop_mean)
print(f'Real Central sample mean: {central_sample.mean():.2f}')
print(f'Real hypothesized mean (H0): {pop_mean:.2f}')
print(f'Real t-statistic: {t_stat:.3f}')
print(f'Real p-value: {p_value:.4f}')
The real p-value here is the probability of seeing a sample mean at least this far from the hypothesized value, if H0 were actually true. It is small, well under the conventional 0.05 threshold, or significance level, alpha. That's evidence against H0, so we reject it: this real sample gives us reason to believe Central really does differ from the overall average, which lines up with what we already confirmed directly in video six using the full real population.
Part 2: What a P-Value Actually Is, Built from Scratch#
office_sample = df[df['Category'] == 'Office Supplies'].sample(n=100, random_state=3)['Sales'].values
tech_sample = df[df['Category'] == 'Technology'].sample(n=100, random_state=3)['Sales'].values
observed_diff = tech_sample.mean() - office_sample.mean()
print(f'Real observed difference (Technology - Office Supplies): {observed_diff:.2f}')
rng = np.random.default_rng(seed=3)
combined = np.concatenate([office_sample, tech_sample])
n1 = len(office_sample)
perm_diffs = np.empty(10000)
for i in range(10000):
shuffled = rng.permutation(combined)
perm_diffs[i] = shuffled[n1:].mean() - shuffled[:n1].mean()
p_perm = (np.abs(perm_diffs) >= abs(observed_diff)).mean()
print(f'Real permutation p-value: {p_perm:.4f}')
import matplotlib.pyplot as plt
plt.figure(figsize=(9, 4))
plt.hist(perm_diffs, bins=50, color='steelblue', edgecolor='white')
plt.axvline(observed_diff, color='darkred', linewidth=2, label=f'Real observed diff = {observed_diff:.1f}')
plt.title('Simulated Null Distribution (Permutation Test) vs. the Real Observed Difference')
plt.xlabel('Difference in Means Under Shuffled Labels')
plt.legend()
plt.show()
t_param, p_param = stats.ttest_ind(tech_sample, office_sample)
print(f'Real parametric t-test p-value: {p_param:.4f}')
print(f'Real permutation p-value: {p_perm:.4f}')
Part 3: Type I and Type II Errors#
rng2 = np.random.default_rng(seed=42)
central_vals = df[df['Region'] == 'Central']['Sales'].values
false_positives = 0
trials = 2000
for _ in range(trials):
a = rng2.choice(central_vals, size=60, replace=False)
b = rng2.choice(central_vals, size=60, replace=False)
_, p = stats.ttest_ind(a, b)
if p < 0.05:
false_positives += 1
print(f'Real Type I error rate (H0 genuinely true here): {false_positives / trials:.1%}')
office_vals = df[df['Category'] == 'Office Supplies']['Sales'].values
tech_vals = df[df['Category'] == 'Technology']['Sales'].values
for n in [10, 30, 80, 200]:
significant = 0
trials2 = 1000
for _ in range(trials2):
a = rng2.choice(office_vals, size=n, replace=False)
b = rng2.choice(tech_vals, size=n, replace=False)
_, p = stats.ttest_ind(a, b)
if p < 0.05:
significant += 1
power = significant / trials2
print(f'n={n}: real statistical power = {power:.1%} (Type II error rate = {1 - power:.1%})')
Contrast this with West versus Central Sales from video six: that real difference is only about $11, tiny next to a standard deviation in the hundreds. Rerunning this exact power simulation on that real pair (try it yourself) shows power barely rising above alpha even at the full real sample size, because the true real effect is nearly swallowed by noise. Not every real, genuine difference is practically detectable, and knowing that distinction before collecting data is exactly what video eight's power calculations are for.
Part 4: One-Sided vs. Two-Sided Tests#
t_two, p_two = stats.ttest_ind(tech_sample, office_sample, alternative='two-sided')
t_one, p_one = stats.ttest_ind(tech_sample, office_sample, alternative='greater')
print(f'Two-sided H1 (Technology != Office Supplies): p = {p_two:.5f}')
print(f"One-sided H1 (Technology > Office Supplies): p = {p_one:.5f}")
Part 5: Common Misinterpretations, Set Straight#
- A p-value is not the probability that H0 is true. It's the probability of seeing data this extreme, assuming H0 already is true.
- A small p-value is not proof of a large or important effect. The West-versus-Central case here has a real, statistically detectable-in-principle difference that's practically tiny; statistical significance and practical significance are separate questions.
- Failing to reject H0 is not proof H0 is true. It might just mean the real sample was too small to detect a genuine effect, exactly what the Type II simulation in part three demonstrated directly.
- Running many tests and reporting only the significant ones inflates the real false-positive rate far above the stated alpha; video ten covers exactly how to correct for that.
Wrap-Up: What You Learned#
- Setting up null and alternative hypotheses, and running a one-sample t-test on real Central region data.
- Building a p-value from scratch with a permutation test, and seeing exactly what 'assuming H0 is true' means as a simulated distribution.
- Type I errors (false alarms) and Type II errors (missed real effects), both measured directly through simulation, including statistical power climbing with sample size.
- One-sided versus two-sided tests, and why the choice must come before looking at the real data.
- Four common p-value misinterpretations, grounded in real examples from this video and the last.
- Video eight puts all of this into a full A/B test: sample size and power calculations planned in advance, run on clearly labeled synthetic experiment data since no real randomized trial exists in this dataset. Subscribe so it lands automatically see you there.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



