Lesson 5 · Statistics for data analysts
Sampling Methods & How Bias Sneaks In | Statistics #5
Video five of the 15-part series: how you draw a sample matters just as much as how big it is. Real Sample Superstore order data, including a real example…
- CourseStatistics for data analysts
- Lesson5 of 15
- Video16 min
- FormatJupyter notebook · 10 code cells
- Data1 dataset
What you'll learn
Datasets used in this lesson
Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.
- superstore_sales.csv715.5 KB
📓 Full notebook
Download .ipynbStatistics for Data Analysts, Video 5: Sampling Methods and Bias#
- Video five of the 15-part series: how you draw a sample matters just as much as how big it is.
- Real Sample Superstore order data, including a real example of sampling bias hiding in plain sight.
- Let's get into it.
Before You Start#
- Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
- You'll need pandas and NumPy.
- Place superstore_sales.csv in the same folder as this notebook.
import pandas as pd
import numpy as np
df = pd.read_csv('superstore_sales.csv')
df['Order Date'] = pd.to_datetime(df['Order Date'])
pop_mean = df['Sales'].mean()
print(f'Real population size: {len(df)}')
print(f'Real population mean Sales: {pop_mean:.2f}')
Part 1: Simple Random Sampling#
srs_sample = df.sample(n=200, random_state=1)
print(f'Real SRS sample mean: {srs_sample["Sales"].mean():.2f}')
print(f'Real population mean: {pop_mean:.2f}')
print(f'Real category mix in sample:')
print(srs_sample['Category'].value_counts(normalize=True).round(3))
Part 2: Stratified Sampling#
def stratified_sample(data, strata_col, total_n, seed):
proportions = data[strata_col].value_counts(normalize=True)
parts = []
for group, prop in proportions.items():
group_n = max(1, round(total_n * prop))
parts.append(data[data[strata_col] == group].sample(n=group_n, random_state=seed))
return pd.concat(parts)
strat_sample = stratified_sample(df, 'Category', 200, seed=1)
print(f'Real stratified sample size: {len(strat_sample)}')
print(strat_sample['Category'].value_counts(normalize=True).round(3))
print(f'Real stratified sample mean: {strat_sample["Sales"].mean():.2f}')
print(f'Real population mean: {pop_mean:.2f}')
Stratified sampling doesn't just protect the overall mean, it protects estimates within each group. If you specifically cared about Technology orders, and Technology is only about 18 percent of the population, a small simple random sample might grab very few Technology rows by chance, while stratified sampling guarantees a reliable real count every time.
rng = np.random.default_rng(seed=5)
tech_counts_srs = []
for _ in range(2000):
s = df.sample(n=50, random_state=rng.integers(0, 1_000_000))
tech_counts_srs.append((s['Category'] == 'Technology').sum())
tech_counts_srs = np.array(tech_counts_srs)
print(f'Real simulated SRS (n=50): mean Technology count = {tech_counts_srs.mean():.2f}, min = {tech_counts_srs.min()}, max = {tech_counts_srs.max()}')
Part 3: Systematic Sampling and a Hidden Trap#
df_sorted = df.sort_values('Order Date').reset_index(drop=True)
k = 8
for offset in range(k):
systematic_sample = df_sorted.iloc[offset::k]
print(f'offset={offset}: n={len(systematic_sample)}, real mean Sales={systematic_sample["Sales"].mean():.2f}')
Every one of those eight samples used the identical systematic method, same step size, same underlying real data, and the resulting real means swing noticeably depending purely on which offset you started from. Because the data is grouped by date with a varying number of order lines per day, a step size close to that daily average causes the sample to consistently land on similar positions within each day's block, quietly injecting real bias that has nothing to do with sample size. The fix is straightforward: either shuffle the real data before sampling systematically, or pick a step size with no relationship to any structure in how the data is ordered.
df_shuffled = df.sample(frac=1, random_state=9).reset_index(drop=True)
for offset in [0, 3, 6]:
systematic_shuffled = df_shuffled.iloc[offset::k]
print(f'shuffled offset={offset}: real mean Sales={systematic_shuffled["Sales"].mean():.2f}')
print(f'Real population mean for reference: {pop_mean:.2f}')
Part 4: A Real Convenience Sampling Trap#
q1_only = df[df['Order Date'].dt.month.isin([1, 2, 3])]
q1_mean = q1_only['Sales'].mean()
pct_off = (q1_mean - pop_mean) / pop_mean * 100
print(f'Real Q1-only convenience sample: n={len(q1_only)}, mean Sales={q1_mean:.2f}')
print(f'Real full population mean Sales: {pop_mean:.2f}')
print(f'Real bias: convenience sample is {pct_off:.1f}% off from the true population mean')
monthly_means = df.groupby(df['Order Date'].dt.month)['Sales'].mean()
print(monthly_means.round(2))
Part 5: Comparing Bias and Variance Across Methods#
srs_means = np.array([df.sample(n=200, random_state=i)['Sales'].mean() for i in range(500)])
strat_means = np.array([stratified_sample(df, 'Category', 200, seed=i)['Sales'].mean() for i in range(500)])
q1_means = np.array([q1_only.sample(n=200, random_state=i)['Sales'].mean() for i in range(500)])
print(f'Real SRS: avg={srs_means.mean():.2f}, std={srs_means.std():.2f}')
print(f'Real stratified: avg={strat_means.mean():.2f}, std={strat_means.std():.2f}')
print(f'Real Q1-only convenience: avg={q1_means.mean():.2f}, std={q1_means.std():.2f}')
print(f'True real population mean: {pop_mean:.2f}')
Wrap-Up: What You Learned#
- Simple random sampling: every real row gets an equal chance, the unbiased baseline.
- Stratified sampling: guarantees proportional representation of real groups, reducing sample-to-sample variance for group-level estimates.
- Systematic sampling: fast, but a real trap when the step size lines up with hidden structure in how the data is ordered.
- Convenience sampling: a real, honest example of structural bias that repetition and bigger samples cannot fix.
- Bias versus variance: SRS and stratified sampling are unbiased on average; a convenience sample is consistently, structurally wrong.
- Video six builds directly on this with confidence intervals, quantifying exactly how much uncertainty a properly drawn sample still carries. Subscribe so it lands automatically see you there.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



