Mathew K Analytics

Lesson 12 · Probability and Statistics in python

Understanding Statistical Sampling and the Law of Large Numbers in Biblical Context

In this lesson, we will explore how probability, sampling, and statistics connect. What is a sample? Why do we collect samples? How does a large sample help…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Introduction: Sampling and the Law of Large Numbers#

In this lesson, we will explore how probability, sampling, and statistics connect.

  • What is a sample?
  • Why do we collect samples?
  • How does a large sample help us understand a population?

You will use Python to see these ideas in action!

import warnings
warnings.filterwarnings("ignore")

# Import core libraries
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt

What is a Sample?#

A sample is a smaller group taken from a larger population.

We almost never see every possible person or event.

By picking a sample, we can estimate what the whole population is like.

# Let us pretend we have a giant jar with 10,000 marbles.
# Some marbles are blue, and some are red.

population_size = 10000
prob_blue = 0.6  # 60% of marbles are blue

population = np.random.choice(['blue', 'red'], size=population_size, p=[prob_blue, 1-prob_blue])

print('Population size:', len(population))
print('First 10 marbles:', population[:10])
Population size: 10000
First 10 marbles: ['blue' 'red' 'red' 'blue' 'red' 'blue' 'blue' 'blue' 'blue' 'red']
# How often is each color in our whole population?
unique, counts = np.unique(population, return_counts=True)
color_counts = dict(zip(unique, counts))
print(color_counts)
{np.str_('blue'): np.int64(6049), np.str_('red'): np.int64(3951)}

Real Life: Why Sample Instead of Measuring All?#

It is usually too expensive or impossible to check every single item in a population. Instead, we use samples to estimate things like

  • Winning vote percentages
  • How tall people are in a city
  • Rate of a disease

The Law of Large Numbers is what lets us trust our sample results.

# Let us take just 20 random marbles from our giant jar.
sample_small = np.random.choice(population, size=20, replace=False)

print('Sample:', sample_small)
unique, counts = np.unique(sample_small, return_counts=True)
sample_counts = dict(zip(unique, counts))
print('Sample color counts:', sample_counts)
Sample: ['blue' 'red' 'blue' 'blue' 'blue' 'blue' 'red' 'blue' 'red' 'red' 'blue'
 'red' 'red' 'blue' 'red' 'red' 'blue' 'red' 'red' 'red']
Sample color counts: {np.str_('blue'): np.int64(9), np.str_('red'): np.int64(11)}
# What if we take bigger samples?
sample_sizes = [20, 50, 200, 1000]

for size in sample_sizes:
    sample = np.random.choice(population, size=size, replace=False)
    prob_blue_sample = np.mean(sample == 'blue')
    print(f'Sample size {size}: fraction blue = {prob_blue_sample:.3f}')
    
Sample size 20: fraction blue = 0.550
Sample size 50: fraction blue = 0.500
Sample size 200: fraction blue = 0.610
Sample size 1000: fraction blue = 0.598

The Law of Large Numbers#

The Law of Large Numbers says: As you increase your sample size, your sample statistics (like average or proportion) get closer and closer to the real values of the whole population.

This is why surveys and experiments use enough data small samples can be noisy!

# Visualize: Let us repeat sampling many times and plot how the blue fraction behaves.
repeats = 300
big_sample_size = 500
fractions = []

for i in range(repeats):
    sample = np.random.choice(population, size=big_sample_size, replace=False)
    blue_fraction = np.mean(sample == 'blue')
    fractions.append(blue_fraction)

plt.hist(fractions, bins=20, color='skyblue', edgecolor='k')
plt.axvline(x=prob_blue, color='red', linestyle='dashed', label='True Blue Proportion')
plt.xlabel('Fraction blue in sample')
plt.ylabel('Count (out of 300 samples)')
plt.title('Sampling distribution: Fraction of blue marbles in large samples')
plt.legend()
plt.show()
No description has been provided for this image
# Practice: Try your own sample size!
user_size = int(input('Enter your sample size (between 5 and 5000): '))

my_sample = np.random.choice(population, size=user_size, replace=False)
fraction_blue = np.mean(my_sample == 'blue')
print(f'Your sample has {fraction_blue:.3f} blue marbles')
Your sample has 0.633 blue marbles

Sampling in Real Datasets#

Let us see a real example using actual data. The Titanic dataset has information on over 800 people from the famous ship.

We will sample from it, just like with our marbles.

# Data setup
import seaborn as sns
titanic = sns.load_dataset('titanic')

print('Shape:', titanic.shape)
titanic.head()
Shape: (891, 15)
survived pclass sex age sibsp parch fare embarked class who adult_male deck embark_town alive alone
0 0 3 male 22.0 1 0 7.2500 S Third man True NaN Southampton no False
1 1 1 female 38.0 1 0 71.2833 C First woman False C Cherbourg yes False
2 1 3 female 26.0 0 0 7.9250 S Third woman False NaN Southampton yes True
3 1 1 female 35.0 1 0 53.1000 S First woman False C Southampton yes False
4 0 3 male 35.0 0 0 8.0500 S Third man True NaN Southampton no True
# Random sampling in Titanic: estimate survival rate from a small group
n = 50
sample_titanic = titanic.sample(n=n, random_state=1)
survival_rate_sample = sample_titanic['survived'].mean()
actual_survival = titanic['survived'].mean()
print(f'Sample size: {n}')
print(f'Sample survival rate: {survival_rate_sample:.2f}')
print(f'Actual dataset survival rate: {actual_survival:.2f}')
Sample size: 50
Sample survival rate: 0.48
Actual dataset survival rate: 0.38
# What if we sample many times?
sample_size = 50
n_repeats = 200
survival_rates = []

for i in range(n_repeats):
    s = titanic.sample(n=sample_size, replace=True)
    survival_rates.append(s['survived'].mean())

plt.hist(survival_rates, bins=15, color='gray', edgecolor='black')
plt.axvline(actual_survival, color='red', linestyle='dotted', label='Actual Survival')
plt.xlabel('Sample Survival Rate')
plt.ylabel('Count')
plt.title('How sample survival rates spread around real value')
plt.legend()
plt.show()
No description has been provided for this image

Mini-Project: Build and Analyze Your Own Sampling Simulation#

Let us put it together!

Your goal: Simulate a population, draw samples, and show how sample statistics move closer to the population value as you increase sample size.

Ready to try? Let us code next.

# Your turn! Simulate a population with two categories.
# For example, say eighty percent apples and twenty percent oranges.
N = 2000
p_apple = 0.8
pop = np.random.choice(['apple', 'orange'], size=N, p=[p_apple, 1 - p_apple])
print('Population size:', N)
Population size: 2000
# Explore: How close do samples get to the apple proportion?
sizes = [5, 20, 100, 500, 1000]
for sz in sizes:
    s = np.random.choice(pop, size=sz, replace=False)
    apple_frac = np.mean(s == 'apple')
    print(f'Sample size {sz}: apple proportion = {apple_frac:.3f}')
    
Sample size 5: apple proportion = 1.000
Sample size 20: apple proportion = 0.900
Sample size 100: apple proportion = 0.820
Sample size 500: apple proportion = 0.812
Sample size 1000: apple proportion = 0.812

Recap: What Did We Learn?#

  • Samples are small groups from larger populations.
  • The Law of Large Numbers explains why bigger samples give more reliable estimates.

Sampling reduces the cost and time needed to learn about a whole population.

Want more practice? Try different population sizes or sample percentages.

Before You Go...#

If this helped you, subscribe and share the video! Keep exploring Data science is all about asking questions.

Try using what you learned on your own data, and see what patterns you find.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.