Lesson 53 · Python for Data Science
3 Sampling and Confidence Intervals in Python
Welcome! Today we will learn how to use sampling and confidence intervals in Python. These are ways to make sense of lots of numbers, and to guess things…
- CoursePython for Data Science
- Lesson53 of 38
- Video14 min
- FormatJupyter notebook · 17 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynb
Sampling and Confidence Intervals in Python#
Welcome! Today we will learn how to use sampling and confidence intervals in Python.
These are ways to make sense of lots of numbers, and to guess things about big groups even if we only look at a few members.
By the end, you will be able to do real-world data sampling and estimate ranges for answers.
# Let us check our Python version
import sys
print('Python version:', sys.version)
What is Sampling?#
Sampling is when we pick a few items from a big group to study.
We do this because it is often too hard or expensive to check every member in a large group.
For example, testing a few pieces of candy from a bag to guess if they all taste good.
# Let us create a fake population list
population = list(range(1, 101)) # Numbers 1 to 100
# Preview the first 10 people in our population
print(population[:10])
Taking a Simple Random Sample#
We can use random sampling to pick a few members from our big group.
This way, every person has a fair chance of being chosen.
Python's random module helps us with this.
import random
# Take a random sample of 10 people
sample = random.sample(population, 10)
print(sample)
# Calculate the mean of the population and the sample
import statistics
pop_mean = statistics.mean(population)
sample_mean = statistics.mean(sample)
print('Population mean:', pop_mean)
print('Sample mean:', sample_mean)
What are Confidence Intervals?#
A confidence interval is a range.
It tells us where we think the true answer is, given our sample.
We are never one hundred percent sure, but we can provide a good guess with a level of confidence, like 95%.
# Calculate the standard error of the mean
import math
sample_std = statistics.stdev(sample)
n = len(sample)
sem = sample_std / math.sqrt(n)
print('Standard error:', sem)
# Build a 95 percent confidence interval for the mean
conf_level = 0.95
z = 1.96 # Z-score for 95 percent confidence
ci_lower = sample_mean - z * sem
ci_upper = sample_mean + z * sem
print('95 percent confidence interval: (', ci_lower, ',', ci_upper, ')')
# Check if the true population mean is inside our confidence interval
inside = ci_lower <= pop_mean <= ci_upper
print('Is population mean in the interval?', inside)
How Do Larger Samples Help?#
When we take larger samples, our estimates usually get better.
The confidence interval gets smaller, meaning we are more certain.
# Try a much bigger sample to see what changes
big_sample = random.sample(population, 50)
big_mean = statistics.mean(big_sample)
big_std = statistics.stdev(big_sample)
big_sem = big_std / math.sqrt(len(big_sample))
big_ci_lower = big_mean - z * big_sem
big_ci_upper = big_mean + z * big_sem
print('Big sample mean:', big_mean)
print('Confidence interval:', (big_ci_lower, big_ci_upper))
# Try changing the sample size interactively
size = int(input('How many people should we sample? '))
custom_sample = random.sample(population, size)
c_mean = statistics.mean(custom_sample)
c_std = statistics.stdev(custom_sample)
c_sem = c_std / math.sqrt(len(custom_sample))
ci_l = c_mean - z * c_sem
ci_u = c_mean + z * c_sem
print('Sample mean:', c_mean)
print('Confidence interval:', (ci_l, ci_u))
# Let us see the effect of repeating our sampling many times
n_trials = 100
successes = 0
for i in range(n_trials):
trial_sample = random.sample(population, 10)
t_mean = statistics.mean(trial_sample)
t_std = statistics.stdev(trial_sample)
t_sem = t_std / math.sqrt(10)
t_lower = t_mean - z * t_sem
t_upper = t_mean + z * t_sem
if t_lower <= pop_mean <= t_upper:
successes += 1
print(f'Out of {n_trials} samples, our interval captured the true mean {successes} times.')
Mini Project: Candy Quality Sampling#
Imagine you work at a candy factory.
You want to know the average weight of candies in a batch, but you can not weigh every single one.
Let us use Python to estimate the true average with a sample and report our confidence about it.
# First, create a list of candy weights as a stand-in for real measurements
random.seed(42)
candy_population = [round(random.gauss(10, 0.5), 2) for _ in range(500)]
print('Sample of 10 candies:', candy_population[:10])
# Sample and estimate mean candy weight with a confidence interval
n = 25
candy_sample = random.sample(candy_population, n)
candy_mean = statistics.mean(candy_sample)
candy_std = statistics.stdev(candy_sample)
candy_sem = candy_std / math.sqrt(n)
candy_ci_lower = candy_mean - z * candy_sem
candy_ci_upper = candy_mean + z * candy_sem
print(f'Estimated average weight: {candy_mean} grams')
print(f'95 percent confidence interval: ({candy_ci_lower}, {candy_ci_upper})')
# Bonus: Ask the user how many candies they want to sample
n_input = int(input('How many candies do you want to sample this time? '))
custom_candy_sample = random.sample(candy_population, n_input)
m = statistics.mean(custom_candy_sample)
s = statistics.stdev(custom_candy_sample)
s_sem = s / math.sqrt(n_input)
ci_lo = m - z * s_sem
ci_hi = m + z * s_sem
print(f'Your estimate: {m} grams')
print(f'Your 95 percent confidence interval: ({ci_lo}, {ci_hi})')
Recap: Key Takeaways#
Sampling helps us guess about big groups without checking everyone.
Confidence intervals show how certain we are about our guess.
Bigger samples give us tighter, more trustworthy intervals.
These ideas are used in science, medicine, business, and more.
Thank you for learning with us!
If you enjoyed this, please like the video, subscribe, and share it with friends. Let us know your questions or how you might use sampling in your own projects!
See you in the next lesson!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



