Lesson 12 · Probability and Statistics in python
Understanding Statistical Sampling and the Law of Large Numbers in Biblical Context
In this lesson, we will explore how probability, sampling, and statistics connect. What is a sample? Why do we collect samples? How does a large sample help…
- CourseProbability and Statistics in python
- Lesson12 of 35
- Video13 min
- FormatJupyter notebook · 12 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbIntroduction: Sampling and the Law of Large Numbers#
In this lesson, we will explore how probability, sampling, and statistics connect.
- What is a sample?
- Why do we collect samples?
- How does a large sample help us understand a population?
You will use Python to see these ideas in action!
import warnings
warnings.filterwarnings("ignore")
# Import core libraries
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
What is a Sample?#
A sample is a smaller group taken from a larger population.
We almost never see every possible person or event.
By picking a sample, we can estimate what the whole population is like.
# Let us pretend we have a giant jar with 10,000 marbles.
# Some marbles are blue, and some are red.
population_size = 10000
prob_blue = 0.6 # 60% of marbles are blue
population = np.random.choice(['blue', 'red'], size=population_size, p=[prob_blue, 1-prob_blue])
print('Population size:', len(population))
print('First 10 marbles:', population[:10])
# How often is each color in our whole population?
unique, counts = np.unique(population, return_counts=True)
color_counts = dict(zip(unique, counts))
print(color_counts)
Real Life: Why Sample Instead of Measuring All?#
It is usually too expensive or impossible to check every single item in a population. Instead, we use samples to estimate things like
- Winning vote percentages
- How tall people are in a city
- Rate of a disease
The Law of Large Numbers is what lets us trust our sample results.
# Let us take just 20 random marbles from our giant jar.
sample_small = np.random.choice(population, size=20, replace=False)
print('Sample:', sample_small)
unique, counts = np.unique(sample_small, return_counts=True)
sample_counts = dict(zip(unique, counts))
print('Sample color counts:', sample_counts)
# What if we take bigger samples?
sample_sizes = [20, 50, 200, 1000]
for size in sample_sizes:
sample = np.random.choice(population, size=size, replace=False)
prob_blue_sample = np.mean(sample == 'blue')
print(f'Sample size {size}: fraction blue = {prob_blue_sample:.3f}')
The Law of Large Numbers#
The Law of Large Numbers says: As you increase your sample size, your sample statistics (like average or proportion) get closer and closer to the real values of the whole population.
This is why surveys and experiments use enough data small samples can be noisy!
# Visualize: Let us repeat sampling many times and plot how the blue fraction behaves.
repeats = 300
big_sample_size = 500
fractions = []
for i in range(repeats):
sample = np.random.choice(population, size=big_sample_size, replace=False)
blue_fraction = np.mean(sample == 'blue')
fractions.append(blue_fraction)
plt.hist(fractions, bins=20, color='skyblue', edgecolor='k')
plt.axvline(x=prob_blue, color='red', linestyle='dashed', label='True Blue Proportion')
plt.xlabel('Fraction blue in sample')
plt.ylabel('Count (out of 300 samples)')
plt.title('Sampling distribution: Fraction of blue marbles in large samples')
plt.legend()
plt.show()
# Practice: Try your own sample size!
user_size = int(input('Enter your sample size (between 5 and 5000): '))
my_sample = np.random.choice(population, size=user_size, replace=False)
fraction_blue = np.mean(my_sample == 'blue')
print(f'Your sample has {fraction_blue:.3f} blue marbles')
Sampling in Real Datasets#
Let us see a real example using actual data. The Titanic dataset has information on over 800 people from the famous ship.
We will sample from it, just like with our marbles.
# Data setup
import seaborn as sns
titanic = sns.load_dataset('titanic')
print('Shape:', titanic.shape)
titanic.head()
# Random sampling in Titanic: estimate survival rate from a small group
n = 50
sample_titanic = titanic.sample(n=n, random_state=1)
survival_rate_sample = sample_titanic['survived'].mean()
actual_survival = titanic['survived'].mean()
print(f'Sample size: {n}')
print(f'Sample survival rate: {survival_rate_sample:.2f}')
print(f'Actual dataset survival rate: {actual_survival:.2f}')
# What if we sample many times?
sample_size = 50
n_repeats = 200
survival_rates = []
for i in range(n_repeats):
s = titanic.sample(n=sample_size, replace=True)
survival_rates.append(s['survived'].mean())
plt.hist(survival_rates, bins=15, color='gray', edgecolor='black')
plt.axvline(actual_survival, color='red', linestyle='dotted', label='Actual Survival')
plt.xlabel('Sample Survival Rate')
plt.ylabel('Count')
plt.title('How sample survival rates spread around real value')
plt.legend()
plt.show()
Mini-Project: Build and Analyze Your Own Sampling Simulation#
Let us put it together!
Your goal: Simulate a population, draw samples, and show how sample statistics move closer to the population value as you increase sample size.
Ready to try? Let us code next.
# Your turn! Simulate a population with two categories.
# For example, say eighty percent apples and twenty percent oranges.
N = 2000
p_apple = 0.8
pop = np.random.choice(['apple', 'orange'], size=N, p=[p_apple, 1 - p_apple])
print('Population size:', N)
# Explore: How close do samples get to the apple proportion?
sizes = [5, 20, 100, 500, 1000]
for sz in sizes:
s = np.random.choice(pop, size=sz, replace=False)
apple_frac = np.mean(s == 'apple')
print(f'Sample size {sz}: apple proportion = {apple_frac:.3f}')
Recap: What Did We Learn?#
- Samples are small groups from larger populations.
- The Law of Large Numbers explains why bigger samples give more reliable estimates.
Sampling reduces the cost and time needed to learn about a whole population.
Want more practice? Try different population sizes or sample percentages.
Before You Go...#
If this helped you, subscribe and share the video! Keep exploring Data science is all about asking questions.
Try using what you learned on your own data, and see what patterns you find.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



