Mathew K Analytics

Lesson 5 · Statistics for data analysts

Sampling Methods & How Bias Sneaks In | Statistics #5

Video five of the 15-part series: how you draw a sample matters just as much as how big it is. Real Sample Superstore order data, including a real example…

What you'll learn

Datasets used in this lesson

Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.

📓 Full notebook

Download .ipynb

Statistics for Data Analysts, Video 5: Sampling Methods and Bias#

  • Video five of the 15-part series: how you draw a sample matters just as much as how big it is.
  • Real Sample Superstore order data, including a real example of sampling bias hiding in plain sight.
  • Let's get into it.

Before You Start#

  • Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
  • You'll need pandas and NumPy.
  • Place superstore_sales.csv in the same folder as this notebook.
import pandas as pd
import numpy as np

df = pd.read_csv('superstore_sales.csv')
df['Order Date'] = pd.to_datetime(df['Order Date'])
pop_mean = df['Sales'].mean()
print(f'Real population size: {len(df)}')
print(f'Real population mean Sales: {pop_mean:.2f}')
Real population size: 9994
Real population mean Sales: 229.86

Part 1: Simple Random Sampling#

srs_sample = df.sample(n=200, random_state=1)
print(f'Real SRS sample mean: {srs_sample["Sales"].mean():.2f}')
print(f'Real population mean: {pop_mean:.2f}')
print(f'Real category mix in sample:')
print(srs_sample['Category'].value_counts(normalize=True).round(3))
Real SRS sample mean: 252.80
Real population mean: 229.86
Real category mix in sample:
Category
Office Supplies    0.590
Furniture          0.215
Technology         0.195
Name: proportion, dtype: float64

Part 2: Stratified Sampling#

def stratified_sample(data, strata_col, total_n, seed):
    proportions = data[strata_col].value_counts(normalize=True)
    parts = []
    for group, prop in proportions.items():
        group_n = max(1, round(total_n * prop))
        parts.append(data[data[strata_col] == group].sample(n=group_n, random_state=seed))
    return pd.concat(parts)

strat_sample = stratified_sample(df, 'Category', 200, seed=1)
print(f'Real stratified sample size: {len(strat_sample)}')
print(strat_sample['Category'].value_counts(normalize=True).round(3))
Real stratified sample size: 200
Category
Office Supplies    0.605
Furniture          0.210
Technology         0.185
Name: proportion, dtype: float64
print(f'Real stratified sample mean: {strat_sample["Sales"].mean():.2f}')
print(f'Real population mean: {pop_mean:.2f}')
Real stratified sample mean: 222.17
Real population mean: 229.86

Stratified sampling doesn't just protect the overall mean, it protects estimates within each group. If you specifically cared about Technology orders, and Technology is only about 18 percent of the population, a small simple random sample might grab very few Technology rows by chance, while stratified sampling guarantees a reliable real count every time.

rng = np.random.default_rng(seed=5)
tech_counts_srs = []
for _ in range(2000):
    s = df.sample(n=50, random_state=rng.integers(0, 1_000_000))
    tech_counts_srs.append((s['Category'] == 'Technology').sum())
tech_counts_srs = np.array(tech_counts_srs)
print(f'Real simulated SRS (n=50): mean Technology count = {tech_counts_srs.mean():.2f}, min = {tech_counts_srs.min()}, max = {tech_counts_srs.max()}')
Real simulated SRS (n=50): mean Technology count = 9.18, min = 2, max = 20

Part 3: Systematic Sampling and a Hidden Trap#

df_sorted = df.sort_values('Order Date').reset_index(drop=True)
k = 8
for offset in range(k):
    systematic_sample = df_sorted.iloc[offset::k]
    print(f'offset={offset}: n={len(systematic_sample)}, real mean Sales={systematic_sample["Sales"].mean():.2f}')
offset=0: n=1250, real mean Sales=244.57
offset=1: n=1250, real mean Sales=212.10
offset=2: n=1249, real mean Sales=210.07
offset=3: n=1249, real mean Sales=218.87
offset=4: n=1249, real mean Sales=264.50
offset=5: n=1249, real mean Sales=264.14
offset=6: n=1249, real mean Sales=226.30
offset=7: n=1249, real mean Sales=198.32

Every one of those eight samples used the identical systematic method, same step size, same underlying real data, and the resulting real means swing noticeably depending purely on which offset you started from. Because the data is grouped by date with a varying number of order lines per day, a step size close to that daily average causes the sample to consistently land on similar positions within each day's block, quietly injecting real bias that has nothing to do with sample size. The fix is straightforward: either shuffle the real data before sampling systematically, or pick a step size with no relationship to any structure in how the data is ordered.

df_shuffled = df.sample(frac=1, random_state=9).reset_index(drop=True)
for offset in [0, 3, 6]:
    systematic_shuffled = df_shuffled.iloc[offset::k]
    print(f'shuffled offset={offset}: real mean Sales={systematic_shuffled["Sales"].mean():.2f}')
print(f'Real population mean for reference: {pop_mean:.2f}')
shuffled offset=0: real mean Sales=231.43
shuffled offset=3: real mean Sales=236.06
shuffled offset=6: real mean Sales=221.31
Real population mean for reference: 229.86

Part 4: A Real Convenience Sampling Trap#

q1_only = df[df['Order Date'].dt.month.isin([1, 2, 3])]
q1_mean = q1_only['Sales'].mean()
pct_off = (q1_mean - pop_mean) / pop_mean * 100
print(f'Real Q1-only convenience sample: n={len(q1_only)}, mean Sales={q1_mean:.2f}')
print(f'Real full population mean Sales: {pop_mean:.2f}')
print(f'Real bias: convenience sample is {pct_off:.1f}% off from the true population mean')
Real Q1-only convenience sample: n=1377, mean Sales=261.21
Real full population mean Sales: 229.86
Real bias: convenience sample is 13.6% off from the true population mean
monthly_means = df.groupby(df['Order Date'].dt.month)['Sales'].mean()
print(monthly_means.round(2))
Order Date
1     249.15
2     199.17
3     294.55
4     206.23
5     210.92
6     213.00
7     207.38
8     225.27
9     222.45
10    244.59
11    239.61
12    231.03
Name: Sales, dtype: float64

Part 5: Comparing Bias and Variance Across Methods#

srs_means = np.array([df.sample(n=200, random_state=i)['Sales'].mean() for i in range(500)])
strat_means = np.array([stratified_sample(df, 'Category', 200, seed=i)['Sales'].mean() for i in range(500)])
q1_means = np.array([q1_only.sample(n=200, random_state=i)['Sales'].mean() for i in range(500)])
print(f'Real SRS: avg={srs_means.mean():.2f}, std={srs_means.std():.2f}')
print(f'Real stratified: avg={strat_means.mean():.2f}, std={strat_means.std():.2f}')
print(f'Real Q1-only convenience: avg={q1_means.mean():.2f}, std={q1_means.std():.2f}')
print(f'True real population mean: {pop_mean:.2f}')
Real SRS: avg=230.93, std=45.42
Real stratified: avg=230.19, std=42.33
Real Q1-only convenience: avg=255.81, std=55.87
True real population mean: 229.86

Wrap-Up: What You Learned#

  • Simple random sampling: every real row gets an equal chance, the unbiased baseline.
  • Stratified sampling: guarantees proportional representation of real groups, reducing sample-to-sample variance for group-level estimates.
  • Systematic sampling: fast, but a real trap when the step size lines up with hidden structure in how the data is ordered.
  • Convenience sampling: a real, honest example of structural bias that repetition and bigger samples cannot fix.
  • Bias versus variance: SRS and stratified sampling are unbiased on average; a convenience sample is consistently, structurally wrong.
  • Video six builds directly on this with confidence intervals, quantifying exactly how much uncertainty a properly drawn sample still carries. Subscribe so it lands automatically see you there.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.