Lesson 13 · Probability and Statistics in python
Understanding the Central Limit Theorem: Key Concepts and Practical Applications
In this hands-on Python session, we will explore the Central Limit Theorem (CLT), see why it is a cornerstone of statistics, and test its power with simple…
- CourseProbability and Statistics in python
- Lesson13 of 35
- Video14 min
- FormatJupyter notebook · 15 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbWelcome to Central Limit Theorem and Its Applications!#
In this hands-on Python session, we will explore the Central Limit Theorem (CLT), see why it is a cornerstone of statistics, and test its power with simple experiments.
Let us break down complex ideas and use real-world data so you see how theory meets practice.
# Suppress annoying warnings to keep things clean
import warnings; warnings.filterwarnings("ignore")
What is the Central Limit Theorem (CLT)?#
CLT says: No matter what your population looks like, if you take lots of samples and average them, their means will look like a bell curve (a normal distribution) as your sample gets big!
This is why averages of things (like average test scores, average heights) tend to follow a normal pattern, even if single measurements are not normal.
# Data setup
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sns
# To make everything look the same each time we run
np.random.seed(42)
# Create a skewed population: a fake set of 10000 people's incomes
population = np.random.exponential(scale=20000, size=10000)
print(f"Population mean: {np.mean(population):.2f}")
print(f"Population std: {np.std(population):.2f}")
# Let us look at this population with a histogram
plt.figure(figsize=(8,4))
sns.histplot(population, bins=50, kde=False, color="skyblue")
plt.title("Histogram of Fake Income Population (not normal)")
plt.xlabel("Income")
plt.ylabel("Count")
plt.show()
What happens if we take just one sample?#
If we picked just a few people, their average might be very far from the true average. Let us see how that works...
# Take a small random sample of 5 incomes
sample = np.random.choice(population, size=5, replace=False)
print("Sample incomes:", sample)
print(f"Sample mean: {np.mean(sample):.2f}")
What does the Central Limit Theorem say?#
It says that as we take more samples of the same size and plot their means, the distribution of those means will look more and more like a normal curve.
# Take 5000 samples of 5, find the mean each time
means_small = [np.mean(np.random.choice(population, size=5, replace=False)) for _ in range(5000)]
plt.figure(figsize=(8,4))
sns.histplot(means_small, bins=40, kde=True, color="orange")
plt.title("Means of 5000 samples (size 5)")
plt.xlabel("Sample Mean")
plt.ylabel("Count")
plt.show()
# Try with bigger samples! Sample size 30 this time
means_bigger = [np.mean(np.random.choice(population, size=30, replace=False)) for _ in range(5000)]
plt.figure(figsize=(8,4))
sns.histplot(means_bigger, bins=40, kde=True, color="green")
plt.title("Means of 5000 samples (size 30)")
plt.xlabel("Sample Mean")
plt.ylabel("Count")
plt.show()
Summary of What We See#
- When you take small samples from a weird population, the averages can jump around.
- When you take more and bigger samples, the means act normal, like a bell.
This is the magic of the central limit theorem!
# Practice: Try your own sample size!
size = int(input("Enter a sample size (try 10, 50, or 100): "))
n_samples = 3000
experiment = [np.mean(np.random.choice(population, size=size, replace=False)) for _ in range(n_samples)]
plt.figure(figsize=(8,4))
sns.histplot(experiment, bins=40, kde=True, color="purple")
plt.title(f"Means of {n_samples} samples (size {size})")
plt.xlabel("Sample Mean")
plt.ylabel("Count")
plt.show()
# Compare means and standard deviations
print(f"Population mean: {np.mean(population):.2f}")
print(f"Sample means average: {np.mean(experiment):.2f}")
print(f"Population std dev: {np.std(population):.2f}")
print(f"Sample means std dev: {np.std(experiment):.2f}")
# Let us apply CLT to a real dataset: the 'tips' dataset
tips = sns.load_dataset("tips")
print("Tips data shape:", tips.shape)
display(tips.head())
# What is the distribution of tips? Let us see if it is normal
plt.figure(figsize=(8,4))
sns.histplot(tips["tip"], bins=30, kde=True, color="salmon")
plt.title("Distribution of Tip Amounts")
plt.xlabel("Tip (dollars)")
plt.ylabel("Count")
plt.show()
# Try the CLT with tips: sample means
tips_means = [tips["tip"].sample(25, replace=False, random_state=i).mean() for i in range(2500)]
plt.figure(figsize=(8,4))
sns.histplot(tips_means, bins=35, kde=True, color="deepskyblue")
plt.title("Means of 2500 samples (size 25) from Tips")
plt.xlabel("Sample Mean Tip")
plt.ylabel("Count")
plt.show()
Best Practices and Troubleshooting#
If your samples seem odd, check these:
- Make sure you are sampling randomly, not in order.
- Use enough samples a few hundred at least.
- If your code is slow, lower the number of samples.
- When in doubt, plot what you are sampling.
Getting weird errors? Try running the cell again, or double-check your brackets and parentheses.
# Extra tip: Visualize how sample size affects normality
fig, axs = plt.subplots(1, 3, figsize=(18,4))
for i, size in enumerate([3, 10, 50]):
means = [np.mean(np.random.choice(population, size=size, replace=False)) for _ in range(2000)]
sns.histplot(means, bins=30, kde=True, ax=axs[i], color="steelblue")
axs[i].set_title(f"Sample size = {size}")
axs[i].set_xlabel("Sample Mean")
axs[i].set_ylabel("Count")
plt.tight_layout()
plt.show()
# Mini-project: Simulate a game (rolling dice) and show the CLT in action
dice_pop = np.random.choice([1,2,3,4,5,6], size=10000, replace=True)
sample_size = 20
means = [np.mean(np.random.choice(dice_pop, size=sample_size, replace=False)) for _ in range(3000)]
plt.figure(figsize=(8,4))
sns.histplot(means, bins=25, kde=True, color="gold")
plt.title("Means of 3000 samples (size 20) from dice rolls")
plt.xlabel("Sample Mean (Dice Roll)")
plt.ylabel("Count")
plt.show()
# Challenge: Use CLT to check average total_bill in 'tips' for random samples
total_bill_means = [tips["total_bill"].sample(40, replace=False, random_state=i).mean() for i in range(1500)]
plt.figure(figsize=(8,4))
sns.histplot(total_bill_means, bins=30, kde=True, color="tomato")
plt.title("Means of 1500 samples (size 40) of Total Bill")
plt.xlabel("Sample Mean of Total Bill")
plt.ylabel("Count")
plt.show()
Recap: Key Takeaways#
- The central limit theorem helps us trust averages.
- No matter the original shape, averages of lots of samples turn bell-like.
- This makes statistics powerful for science, business, and understanding daily life.
Keep practicing sampling and watch those bell curves appear!
Thank you for learning with us!#
If you enjoyed this notebook or video, give it a thumbs up and subscribe. Keep exploring, and let the data surprise you!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



