Mathew K Analytics

Lesson 13 · Probability and Statistics in python

Understanding the Central Limit Theorem: Key Concepts and Practical Applications

In this hands-on Python session, we will explore the Central Limit Theorem (CLT), see why it is a cornerstone of statistics, and test its power with simple…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Welcome to Central Limit Theorem and Its Applications!#

In this hands-on Python session, we will explore the Central Limit Theorem (CLT), see why it is a cornerstone of statistics, and test its power with simple experiments.

Let us break down complex ideas and use real-world data so you see how theory meets practice.

# Suppress annoying warnings to keep things clean
import warnings; warnings.filterwarnings("ignore")

What is the Central Limit Theorem (CLT)?#

CLT says: No matter what your population looks like, if you take lots of samples and average them, their means will look like a bell curve (a normal distribution) as your sample gets big!

This is why averages of things (like average test scores, average heights) tend to follow a normal pattern, even if single measurements are not normal.

# Data setup
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sns

# To make everything look the same each time we run
np.random.seed(42)
# Create a skewed population: a fake set of 10000 people's incomes
population = np.random.exponential(scale=20000, size=10000)
print(f"Population mean: {np.mean(population):.2f}")
print(f"Population std: {np.std(population):.2f}")
Population mean: 19549.98
Population std: 19487.12
# Let us look at this population with a histogram
plt.figure(figsize=(8,4))
sns.histplot(population, bins=50, kde=False, color="skyblue")
plt.title("Histogram of Fake Income Population (not normal)")
plt.xlabel("Income")
plt.ylabel("Count")
plt.show()
No description has been provided for this image

What happens if we take just one sample?#

If we picked just a few people, their average might be very far from the true average. Let us see how that works...

# Take a small random sample of 5 incomes
sample = np.random.choice(population, size=5, replace=False)
print("Sample incomes:", sample)
print(f"Sample mean: {np.mean(sample):.2f}")
Sample incomes: [27997.98460233 11324.87199283 31971.30952896 35234.21990535
 50688.17434222]
Sample mean: 31443.31

What does the Central Limit Theorem say?#

It says that as we take more samples of the same size and plot their means, the distribution of those means will look more and more like a normal curve.

# Take 5000 samples of 5, find the mean each time
means_small = [np.mean(np.random.choice(population, size=5, replace=False)) for _ in range(5000)]

plt.figure(figsize=(8,4))
sns.histplot(means_small, bins=40, kde=True, color="orange")
plt.title("Means of 5000 samples (size 5)")
plt.xlabel("Sample Mean")
plt.ylabel("Count")
plt.show()
No description has been provided for this image
# Try with bigger samples! Sample size 30 this time
means_bigger = [np.mean(np.random.choice(population, size=30, replace=False)) for _ in range(5000)]

plt.figure(figsize=(8,4))
sns.histplot(means_bigger, bins=40, kde=True, color="green")
plt.title("Means of 5000 samples (size 30)")
plt.xlabel("Sample Mean")
plt.ylabel("Count")
plt.show()
No description has been provided for this image

Summary of What We See#

  • When you take small samples from a weird population, the averages can jump around.
  • When you take more and bigger samples, the means act normal, like a bell.

This is the magic of the central limit theorem!

# Practice: Try your own sample size!
size = int(input("Enter a sample size (try 10, 50, or 100): "))
n_samples = 3000
experiment = [np.mean(np.random.choice(population, size=size, replace=False)) for _ in range(n_samples)]

plt.figure(figsize=(8,4))
sns.histplot(experiment, bins=40, kde=True, color="purple")
plt.title(f"Means of {n_samples} samples (size {size})")
plt.xlabel("Sample Mean")
plt.ylabel("Count")
plt.show()
No description has been provided for this image
# Compare means and standard deviations
print(f"Population mean: {np.mean(population):.2f}")
print(f"Sample means average: {np.mean(experiment):.2f}")
print(f"Population std dev: {np.std(population):.2f}")
print(f"Sample means std dev: {np.std(experiment):.2f}")
Population mean: 19549.98
Sample means average: 19569.41
Population std dev: 19487.12
Sample means std dev: 2755.03
# Let us apply CLT to a real dataset: the 'tips' dataset
tips = sns.load_dataset("tips")
print("Tips data shape:", tips.shape)
display(tips.head())
Tips data shape: (244, 7)
total_bill tip sex smoker day time size
0 16.99 1.01 Female No Sun Dinner 2
1 10.34 1.66 Male No Sun Dinner 3
2 21.01 3.50 Male No Sun Dinner 3
3 23.68 3.31 Male No Sun Dinner 2
4 24.59 3.61 Female No Sun Dinner 4
# What is the distribution of tips? Let us see if it is normal
plt.figure(figsize=(8,4))
sns.histplot(tips["tip"], bins=30, kde=True, color="salmon")
plt.title("Distribution of Tip Amounts")
plt.xlabel("Tip (dollars)")
plt.ylabel("Count")
plt.show()
No description has been provided for this image
# Try the CLT with tips: sample means
tips_means = [tips["tip"].sample(25, replace=False, random_state=i).mean() for i in range(2500)]
plt.figure(figsize=(8,4))
sns.histplot(tips_means, bins=35, kde=True, color="deepskyblue")
plt.title("Means of 2500 samples (size 25) from Tips")
plt.xlabel("Sample Mean Tip")
plt.ylabel("Count")
plt.show()
No description has been provided for this image

Best Practices and Troubleshooting#

If your samples seem odd, check these:

  • Make sure you are sampling randomly, not in order.
  • Use enough samples a few hundred at least.
  • If your code is slow, lower the number of samples.
  • When in doubt, plot what you are sampling.

Getting weird errors? Try running the cell again, or double-check your brackets and parentheses.

# Extra tip: Visualize how sample size affects normality
fig, axs = plt.subplots(1, 3, figsize=(18,4))
for i, size in enumerate([3, 10, 50]):
    means = [np.mean(np.random.choice(population, size=size, replace=False)) for _ in range(2000)]
    sns.histplot(means, bins=30, kde=True, ax=axs[i], color="steelblue")
    axs[i].set_title(f"Sample size = {size}")
    axs[i].set_xlabel("Sample Mean")
    axs[i].set_ylabel("Count")
plt.tight_layout()
plt.show()
No description has been provided for this image
# Mini-project: Simulate a game (rolling dice) and show the CLT in action
dice_pop = np.random.choice([1,2,3,4,5,6], size=10000, replace=True)
sample_size = 20
means = [np.mean(np.random.choice(dice_pop, size=sample_size, replace=False)) for _ in range(3000)]
plt.figure(figsize=(8,4))
sns.histplot(means, bins=25, kde=True, color="gold")
plt.title("Means of 3000 samples (size 20) from dice rolls")
plt.xlabel("Sample Mean (Dice Roll)")
plt.ylabel("Count")
plt.show()
No description has been provided for this image
# Challenge: Use CLT to check average total_bill in 'tips' for random samples
total_bill_means = [tips["total_bill"].sample(40, replace=False, random_state=i).mean() for i in range(1500)]
plt.figure(figsize=(8,4))
sns.histplot(total_bill_means, bins=30, kde=True, color="tomato")
plt.title("Means of 1500 samples (size 40) of Total Bill")
plt.xlabel("Sample Mean of Total Bill")
plt.ylabel("Count")
plt.show()
No description has been provided for this image

Recap: Key Takeaways#

  • The central limit theorem helps us trust averages.
  • No matter the original shape, averages of lots of samples turn bell-like.
  • This makes statistics powerful for science, business, and understanding daily life.

Keep practicing sampling and watch those bell curves appear!

Thank you for learning with us!#

If you enjoyed this notebook or video, give it a thumbs up and subscribe. Keep exploring, and let the data surprise you!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.