Mathew K Analytics

Lesson 4 · Statistics for data analysts

The Central Limit Theorem, Finally Explained | Statistics #4

Video four of the 15-part series: arguably the single most important idea in all of inferential statistics. Real, heavily skewed Sample Superstore sales and…

What you'll learn

Datasets used in this lesson

Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.

📓 Full notebook

Download .ipynb

Statistics for Data Analysts, Video 4: The Central Limit Theorem#

  • Video four of the 15-part series: arguably the single most important idea in all of inferential statistics.
  • Real, heavily skewed Sample Superstore sales and profit data as the population we sample from.
  • Let's get into it.

Before You Start#

  • Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
  • You'll need pandas, NumPy, and Matplotlib.
  • Place superstore_sales.csv in the same folder as this notebook.
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt

df = pd.read_csv('superstore_sales.csv')
population = df['Sales']
print(f'Real population size: {len(population)}')
Real population size: 9994

Part 1: A Genuinely Non-Normal Population#

pop_mean = population.mean()
pop_std = population.std()
from scipy import stats
pop_skew = stats.skew(population)
print(f'Real population mean: {pop_mean:.2f}')
print(f'Real population std dev: {pop_std:.2f}')
print(f'Real population skewness: {pop_skew:.2f}')
Real population mean: 229.86
Real population std dev: 623.25
Real population skewness: 12.97
plt.figure(figsize=(9, 4))
plt.hist(population, bins=60, color='steelblue', edgecolor='white')
plt.axvline(pop_mean, color='darkred', linewidth=2, label=f'Real population mean = {pop_mean:.0f}')
plt.title('Real Population Distribution: Individual Order Sales')
plt.xlabel('Sales ($)')
plt.legend()
plt.show()
No description has been provided for this image

Part 2: One Small Sample#

rng = np.random.default_rng(seed=3)
sample_30 = rng.choice(population, size=30, replace=False)
print(f'Real sample mean (n=30): {sample_30.mean():.2f}')
print(f'Real population mean: {pop_mean:.2f}')
print(f'Difference: {sample_30.mean() - pop_mean:.2f}')
Real sample mean (n=30): 348.91
Real population mean: 229.86
Difference: 119.05

One sample gives one estimate, and that estimate has sampling error baked in. The Central Limit Theorem isn't about any single sample, it's about what happens to the distribution of sample means if we could repeat this process over and over.

Part 3: Repeated Sampling - Building the Sampling Distribution#

n_samples = 5000
sample_size = 30
sample_means = np.array([rng.choice(population, size=sample_size, replace=True).mean() for _ in range(n_samples)])
print(f'Simulated {n_samples} samples of size {sample_size} drawn from the real population')
print(f'Mean of the simulated sample means: {sample_means.mean():.2f}')
print(f'Real population mean: {pop_mean:.2f}')
Simulated 5000 samples of size 30 drawn from the real population
Mean of the simulated sample means: 230.33
Real population mean: 229.86
plt.figure(figsize=(9, 4))
plt.hist(sample_means, bins=50, color='steelblue', edgecolor='white', density=True)
x = np.linspace(sample_means.min(), sample_means.max(), 200)
fitted_normal = stats.norm.pdf(x, sample_means.mean(), sample_means.std())
plt.plot(x, fitted_normal, color='darkred', linewidth=2, label='Fitted normal curve')
plt.title('Sampling Distribution of the Mean (n=30, 5,000 simulated samples)')
plt.xlabel('Sample Mean Sales ($)')
plt.legend()
plt.show()
No description has been provided for this image

Part 4: Sample Size and the Standard Error#

fig, axes = plt.subplots(1, 3, figsize=(13, 4), sharex=True)
for ax, n in zip(axes, [5, 30, 200]):
    means_n = np.array([rng.choice(population, size=n, replace=True).mean() for _ in range(3000)])
    ax.hist(means_n, bins=40, color='steelblue', edgecolor='white', density=True)
    ax.set_title(f'n={n}, SE={means_n.std():.1f}')
plt.tight_layout()
plt.show()
No description has been provided for this image
se_formula_30 = pop_std / np.sqrt(30)
se_formula_200 = pop_std / np.sqrt(200)
print(f'Theoretical SE at n=30 (pop_std / sqrt(30)): {se_formula_30:.2f}')
print(f'Simulated SE at n=30 (from part 3): {sample_means.std():.2f}')
print(f'Theoretical SE at n=200 (pop_std / sqrt(200)): {se_formula_200:.2f}')
Theoretical SE at n=30 (pop_std / sqrt(30)): 113.79
Simulated SE at n=30 (from part 3): 114.06
Theoretical SE at n=200 (pop_std / sqrt(200)): 44.07

Part 5: Confirming It Holds for Profit Too#

profit_population = df['Profit']
print(f'Real profit population skewness: {stats.skew(profit_population):.2f}')
print(f'Real profit population min: {profit_population.min():.2f}, max: {profit_population.max():.2f}')
profit_sample_means = np.array([rng.choice(profit_population, size=30, replace=True).mean() for _ in range(5000)])
plt.figure(figsize=(9, 4))
plt.hist(profit_sample_means, bins=50, color='steelblue', edgecolor='white', density=True)
xp = np.linspace(profit_sample_means.min(), profit_sample_means.max(), 200)
plt.plot(xp, stats.norm.pdf(xp, profit_sample_means.mean(), profit_sample_means.std()), color='darkred', linewidth=2)
plt.title('Sampling Distribution of Mean Profit (n=30, real data)')
plt.xlabel('Sample Mean Profit ($)')
plt.show()
Real profit population skewness: 7.56
Real profit population min: -6599.98, max: 8399.98
No description has been provided for this image

Wrap-Up: What You Learned#

  • The real Sales and Profit populations are both heavily skewed, confirmed with real histograms and skewness values.
  • Simulating repeated real sampling shows the distribution of sample means centers on the true population mean and looks approximately normal, even though the population itself doesn't.
  • The standard error shrinks with the square root of sample size, verified against the theoretical formula.
  • This held for two different, genuinely non-normal real variables, not just a cherry-picked one.
  • This is why so many statistical tools assume normality of a mean or an average, even on skewed real data: the Central Limit Theorem is doing the work behind the scenes. Video five turns to sampling methods and the kinds of real bias that can quietly undermine all of this. Subscribe so it lands automatically see you there.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.