Mathew K Analytics

Lesson 6 · Statistics for data analysts

Confidence Intervals Done Properly | Statistics #6

Video six of the 15-part series: not just how to compute a confidence interval, but what it actually means, and an honest look at when it doesn't behave the…

What you'll learn

Datasets used in this lesson

Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.

📓 Full notebook

Download .ipynb

Statistics for Data Analysts, Video 6: Confidence Intervals, Properly#

  • Video six of the 15-part series: not just how to compute a confidence interval, but what it actually means, and an honest look at when it doesn't behave the way the textbook formula promises.
  • Real Sample Superstore order data throughout.
  • Let's get into it.

Before You Start#

  • Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
  • You'll need pandas, NumPy, and SciPy.
  • Place superstore_sales.csv in the same folder as this notebook.
import pandas as pd
import numpy as np
from scipy import stats

df = pd.read_csv('superstore_sales.csv')
pop_mean = df['Sales'].mean()
print(f'Real population mean Sales: {pop_mean:.2f}')
Real population mean Sales: 229.86

Part 1: A Confidence Interval for a Mean#

sample = df.sample(n=100, random_state=21)['Sales']
n = len(sample)
mean = sample.mean()
se = sample.std(ddof=1) / np.sqrt(n)
print(f'Real sample mean: {mean:.2f}')
print(f'Real standard error: {se:.2f}')
Real sample mean: 291.06
Real standard error: 57.47
t_crit = stats.t.ppf(0.975, df=n - 1)
margin = t_crit * se
ci_low, ci_high = mean - margin, mean + margin
print(f'Real t critical value (95% CI, df={n - 1}): {t_crit:.4f}')
print(f'Real 95% CI: ({ci_low:.2f}, {ci_high:.2f})')
print(f'Real true population mean: {pop_mean:.2f}')
print(f'Does this real interval contain the true mean? {ci_low <= pop_mean <= ci_high}')
Real t critical value (95% CI, df=99): 1.9842
Real 95% CI: (177.03, 405.09)
Real true population mean: 229.86
Does this real interval contain the true mean? True

Here's the part that trips people up: a 95% confidence interval does not mean there's a 95% probability the true mean falls in this specific range. The true mean either is or isn't in this real interval, full stop. What 95% confidence actually means is about the procedure: if you repeated this real sampling and interval-building process over and over, about 95% of the resulting intervals would contain the true value. We'll prove that directly in part five.

ci_scipy = stats.t.interval(0.95, df=n - 1, loc=mean, scale=se)
print(f'scipy shortcut, same real result: ({ci_scipy[0]:.2f}, {ci_scipy[1]:.2f})')
scipy shortcut, same real result: (177.03, 405.09)

Part 2: A Confidence Interval for a Proportion#

sample_p = df.sample(n=300, random_state=21)
p_hat = (sample_p['Category'] == 'Technology').mean()
n_p = len(sample_p)
se_p = np.sqrt(p_hat * (1 - p_hat) / n_p)
z_crit = stats.norm.ppf(0.975)
ci_p_low, ci_p_high = p_hat - z_crit * se_p, p_hat + z_crit * se_p
true_p = (df['Category'] == 'Technology').mean()
print(f'Real sample proportion Technology: {p_hat:.4f}')
print(f'Real 95% CI: ({ci_p_low:.4f}, {ci_p_high:.4f})')
print(f'Real true population proportion: {true_p:.4f}')
Real sample proportion Technology: 0.1833
Real 95% CI: (0.1395, 0.2271)
Real true population proportion: 0.1848

Part 3: A Confidence Interval for a Difference of Means#

west = df[df['Region'] == 'West'].sample(n=150, random_state=5)['Sales']
central = df[df['Region'] == 'Central'].sample(n=150, random_state=5)['Sales']
diff = west.mean() - central.mean()
se_diff = np.sqrt(west.var(ddof=1) / 150 + central.var(ddof=1) / 150)
ci_diff_low, ci_diff_high = diff - z_crit * se_diff, diff + z_crit * se_diff
print(f'Real West sample mean: {west.mean():.2f}, Central sample mean: {central.mean():.2f}')
print(f'Real observed difference: {diff:.2f}')
print(f'Real 95% CI for the difference: ({ci_diff_low:.2f}, {ci_diff_high:.2f})')
Real West sample mean: 288.00, Central sample mean: 138.20
Real observed difference: 149.80
Real 95% CI for the difference: (49.75, 249.85)
true_west_mean = df[df['Region'] == 'West']['Sales'].mean()
true_central_mean = df[df['Region'] == 'Central']['Sales'].mean()
true_diff = true_west_mean - true_central_mean
print(f'Real true West population mean: {true_west_mean:.2f}')
print(f'Real true Central population mean: {true_central_mean:.2f}')
print(f'Real true difference: {true_diff:.2f}')
print(f'Does this real interval contain the true difference? {ci_diff_low <= true_diff <= ci_diff_high}')
Real true West population mean: 226.49
Real true Central population mean: 215.77
Real true difference: 10.72
Does this real interval contain the true difference? False

Part 4: Bootstrap Confidence Intervals#

rng = np.random.default_rng(seed=21)
sample_vals = df.sample(n=100, random_state=21)['Sales'].values
sample_median = np.median(sample_vals)
boot_medians = np.array([np.median(rng.choice(sample_vals, size=len(sample_vals), replace=True)) for _ in range(5000)])
boot_ci_low, boot_ci_high = np.percentile(boot_medians, [2.5, 97.5])
print(f'Real sample median: {sample_median:.2f}')
print(f'Real bootstrap 95% CI: ({boot_ci_low:.2f}, {boot_ci_high:.2f})')
Real sample median: 84.01
Real bootstrap 95% CI: (55.26, 120.33)
true_median = df['Sales'].median()
print(f'Real true population median: {true_median:.2f}')
print(f'Does the real bootstrap interval contain the true median? {boot_ci_low <= true_median <= boot_ci_high}')
Real true population median: 54.49
Does the real bootstrap interval contain the true median? False

Part 5: Proving What 95% Confidence Actually Means#

rng2 = np.random.default_rng(seed=100)
sales_values = df['Sales'].values
for sample_n in [30, 100, 500, 1000]:
    hits = 0
    trials = 1000
    for _ in range(trials):
        s = rng2.choice(sales_values, size=sample_n, replace=False)
        m = s.mean()
        se_s = s.std(ddof=1) / np.sqrt(sample_n)
        tcrit_s = stats.t.ppf(0.975, df=sample_n - 1)
        lo, hi = m - tcrit_s * se_s, m + tcrit_s * se_s
        if lo <= pop_mean <= hi:
            hits += 1
    print(f'n={sample_n}: real empirical coverage = {hits / trials:.1%} (target: 95%)')
n=30: real empirical coverage = 79.7% (target: 95%)
n=100: real empirical coverage = 87.1% (target: 95%)
n=500: real empirical coverage = 91.1% (target: 95%)
n=1000: real empirical coverage = 94.2% (target: 95%)

This is the honest, important finding: at small and moderate real sample sizes, the actual coverage runs noticeably below the nominal 95%, because the extreme real skew in Sales, skewness near 13 as we found in video four, means the Central Limit Theorem needs a genuinely large n before the sampling distribution is close enough to normal for this formula to behave as advertised. Coverage climbs steadily toward 95% as n grows, exactly as the theory predicts, it just takes real, heavily skewed data longer to get there than a tidy textbook example would suggest.

Wrap-Up: What You Learned#

  • Confidence intervals for a mean, using the t-distribution, and the correct long-run interpretation of what 'confidence' means.
  • Confidence intervals for a proportion, using the normal approximation.
  • Confidence intervals for a difference of means, including a real, honest miss caused by sample-to-sample variability.
  • Bootstrap confidence intervals, for statistics like the median that don't have a tidy formula.
  • A direct simulation proving that nominal 95% coverage is a real property of the long-run procedure, but that it can run noticeably below target at realistic sample sizes when the underlying real data is heavily skewed.
  • Video seven turns this same machinery toward formal hypothesis testing: null and alternative hypotheses, Type I and Type II errors, and what a p-value actually is. Subscribe so it lands automatically see you there.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.