Lesson 1 · Statistics for data analysts
Probability Foundations for Data Analysts | Statistics #1
This is video one of a 15-part series that goes deeper into the statistics behind everything you've plotted, grouped, and summarized so far. Real Sample…
- CourseStatistics for data analysts
- Lesson1 of 15
- Video17 min
- FormatJupyter notebook · 13 code cells
- Data1 dataset
What you'll learn
Datasets used in this lesson
Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.
- superstore_sales.csv715.5 KB
📓 Full notebook
Download .ipynbStatistics for Data Analysts, Video 1: Probability Foundations#
- This is video one of a 15-part series that goes deeper into the statistics behind everything you've plotted, grouped, and summarized so far.
- Real Sample Superstore order data throughout, plus NumPy simulations you can run yourself.
- We start at the foundation: what probability actually means, and the rules that govern how it combines.
Before You Start#
- Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
- You'll need pandas and NumPy, both already installed if you followed the earlier series.
- Place superstore_sales.csv in the same folder as this notebook.
import pandas as pd
import numpy as np
df = pd.read_csv('superstore_sales.csv')
print(df.shape)
print(df.head(3))
Part 1: What Probability Actually Means#
total_orders = len(df)
tech_orders = (df['Category'] == 'Technology').sum()
p_tech = tech_orders / total_orders
print(f'Real orders: {total_orders}')
print(f'Real Technology orders: {tech_orders}')
print(f'P(Technology) = {p_tech:.4f}')
This is called empirical probability: an estimate built directly from observed real frequencies, as opposed to a theoretical probability you'd derive from first principles, like a fair coin having exactly a 0.5 chance of heads. Most of the probabilities you'll compute as a data analyst are empirical, because you're working with real logged data, not idealized dice.
category_probs = df['Category'].value_counts(normalize=True)
print(category_probs)
print(f'Real probabilities sum to: {category_probs.sum():.4f}')
Part 2: The Rules That Govern Probability#
p_not_tech = 1 - p_tech
check = (df['Category'] != 'Technology').mean()
print(f'P(not Technology) via complement rule: {p_not_tech:.4f}')
print(f'P(not Technology) computed directly: {check:.4f}')
p_furniture = (df['Category'] == 'Furniture').mean()
p_tech_or_furniture = p_tech + p_furniture
check_or = df['Category'].isin(['Technology', 'Furniture']).mean()
print(f'P(Technology) + P(Furniture) = {p_tech_or_furniture:.4f}')
print(f'P(Technology or Furniture) direct = {check_or:.4f}')
Mutual exclusivity matters here: Category and Region are not mutually exclusive, an order has one of each at the same time. When two events can happen together, straight addition double-counts the overlap, and you need the full addition rule: P(A or B) = P(A) + P(B) - P(A and B).
p_tech2 = (df['Category'] == 'Technology').mean()
p_west = (df['Region'] == 'West').mean()
p_tech_and_west = ((df['Category'] == 'Technology') & (df['Region'] == 'West')).mean()
p_tech_or_west = p_tech2 + p_west - p_tech_and_west
check_or2 = ((df['Category'] == 'Technology') | (df['Region'] == 'West')).mean()
print(f'P(Technology or West) via full rule: {p_tech_or_west:.4f}')
print(f'P(Technology or West) direct: {check_or2:.4f}')
Part 3: Independence and Conditional Probability#
p_tech_given_west = df.loc[df['Region'] == 'West', 'Category'].eq('Technology').mean()
p_tech_overall = (df['Category'] == 'Technology').mean()
print(f'P(Technology | West) = {p_tech_given_west:.4f}')
print(f'P(Technology) overall = {p_tech_overall:.4f}')
p_a_and_b = p_tech_and_west
p_b = p_west
p_a_given_b = p_a_and_b / p_b
print(f'P(Technology | West) via definition: {p_a_given_b:.4f}')
print(f'Matches direct calculation: {abs(p_a_given_b - p_tech_given_west) < 1e-9}')
region_category = pd.crosstab(df['Region'], df['Category'], normalize='index')
print(region_category.round(4))
Part 4: Random Variables and Simulating with NumPy#
rng = np.random.default_rng(seed=42)
coin_flips = rng.integers(0, 2, size=10000)
p_heads = coin_flips.mean()
print(f'Simulated P(heads) over 10,000 flips: {p_heads:.4f}')
sampled_orders = rng.choice(df['Category'], size=5000, replace=True)
sampled_probs = pd.Series(sampled_orders).value_counts(normalize=True)
print('Simulated by resampling from the real category distribution:')
print(sampled_probs)
print()
print('Original real distribution:')
print(category_probs)
This resampling technique, drawing new simulated samples from real observed data, is called the bootstrap. It comes back in a much bigger way in video six, when we build confidence intervals without relying on any theoretical formula.
Part 5: The Law of Large Numbers#
import matplotlib.pyplot as plt
sample_sizes = np.arange(10, 5001, 10)
running_flips = rng.integers(0, 2, size=5000)
running_means = np.cumsum(running_flips) / np.arange(1, 5001)
plt.figure(figsize=(9, 4))
plt.plot(running_means, color='steelblue', linewidth=1)
plt.axhline(0.5, color='darkred', linestyle='--', label='True probability (0.5)')
plt.title('Simulated Coin Flip Probability Converging with Sample Size')
plt.xlabel('Number of Flips')
plt.ylabel('Running Probability of Heads')
plt.legend()
plt.show()
for n in [10, 100, 1000, 5000]:
est = running_means[n - 1]
print(f'After {n} flips: estimated P(heads) = {est:.4f}, error = {abs(est - 0.5):.4f}')
Wrap-Up: What You Learned#
- Empirical probability: computing real probabilities directly from observed data.
- The complement, addition, and multiplication-style rules, each verified two different ways on real data.
- Conditional probability, and how to check whether two real variables look independent.
- Simulating random processes with NumPy, including resampling real data with the bootstrap technique.
- The law of large numbers, watched directly as a simulated estimate converges to the truth.
- Video two moves from these general rules to formal discrete distributions: the binomial and the Poisson, applied to real daily order counts. Subscribe so it lands automatically see you there.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



