Mathew K Analytics

Lesson 1 · Statistics for data analysts

Probability Foundations for Data Analysts | Statistics #1

This is video one of a 15-part series that goes deeper into the statistics behind everything you've plotted, grouped, and summarized so far. Real Sample…

What you'll learn

Datasets used in this lesson

Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.

📓 Full notebook

Download .ipynb

Statistics for Data Analysts, Video 1: Probability Foundations#

  • This is video one of a 15-part series that goes deeper into the statistics behind everything you've plotted, grouped, and summarized so far.
  • Real Sample Superstore order data throughout, plus NumPy simulations you can run yourself.
  • We start at the foundation: what probability actually means, and the rules that govern how it combines.

Before You Start#

  • Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
  • You'll need pandas and NumPy, both already installed if you followed the earlier series.
  • Place superstore_sales.csv in the same folder as this notebook.
import pandas as pd
import numpy as np

df = pd.read_csv('superstore_sales.csv')
print(df.shape)
print(df.head(3))
(9994, 9)
  Order Date Region       State         Category Sub-Category    Segment  \
0  11/8/2016  South    Kentucky        Furniture    Bookcases   Consumer   
1  11/8/2016  South    Kentucky        Furniture       Chairs   Consumer   
2  6/12/2016   West  California  Office Supplies       Labels  Corporate   

    Sales    Profit  Quantity  
0  261.96   41.9136         2  
1  731.94  219.5820         3  
2   14.62    6.8714         2  

Part 1: What Probability Actually Means#

total_orders = len(df)
tech_orders = (df['Category'] == 'Technology').sum()
p_tech = tech_orders / total_orders
print(f'Real orders: {total_orders}')
print(f'Real Technology orders: {tech_orders}')
print(f'P(Technology) = {p_tech:.4f}')
Real orders: 9994
Real Technology orders: 1847
P(Technology) = 0.1848

This is called empirical probability: an estimate built directly from observed real frequencies, as opposed to a theoretical probability you'd derive from first principles, like a fair coin having exactly a 0.5 chance of heads. Most of the probabilities you'll compute as a data analyst are empirical, because you're working with real logged data, not idealized dice.

category_probs = df['Category'].value_counts(normalize=True)
print(category_probs)
print(f'Real probabilities sum to: {category_probs.sum():.4f}')
Category
Office Supplies    0.602962
Furniture          0.212227
Technology         0.184811
Name: proportion, dtype: float64
Real probabilities sum to: 1.0000

Part 2: The Rules That Govern Probability#

p_not_tech = 1 - p_tech
check = (df['Category'] != 'Technology').mean()
print(f'P(not Technology) via complement rule: {p_not_tech:.4f}')
print(f'P(not Technology) computed directly: {check:.4f}')
P(not Technology) via complement rule: 0.8152
P(not Technology) computed directly: 0.8152
p_furniture = (df['Category'] == 'Furniture').mean()
p_tech_or_furniture = p_tech + p_furniture
check_or = df['Category'].isin(['Technology', 'Furniture']).mean()
print(f'P(Technology) + P(Furniture) = {p_tech_or_furniture:.4f}')
print(f'P(Technology or Furniture) direct = {check_or:.4f}')
P(Technology) + P(Furniture) = 0.3970
P(Technology or Furniture) direct = 0.3970

Mutual exclusivity matters here: Category and Region are not mutually exclusive, an order has one of each at the same time. When two events can happen together, straight addition double-counts the overlap, and you need the full addition rule: P(A or B) = P(A) + P(B) - P(A and B).

p_tech2 = (df['Category'] == 'Technology').mean()
p_west = (df['Region'] == 'West').mean()
p_tech_and_west = ((df['Category'] == 'Technology') & (df['Region'] == 'West')).mean()
p_tech_or_west = p_tech2 + p_west - p_tech_and_west
check_or2 = ((df['Category'] == 'Technology') | (df['Region'] == 'West')).mean()
print(f'P(Technology or West) via full rule: {p_tech_or_west:.4f}')
print(f'P(Technology or West) direct: {check_or2:.4f}')
P(Technology or West) via full rule: 0.4454
P(Technology or West) direct: 0.4454

Part 3: Independence and Conditional Probability#

p_tech_given_west = df.loc[df['Region'] == 'West', 'Category'].eq('Technology').mean()
p_tech_overall = (df['Category'] == 'Technology').mean()
print(f'P(Technology | West) = {p_tech_given_west:.4f}')
print(f'P(Technology) overall = {p_tech_overall:.4f}')
P(Technology | West) = 0.1870
P(Technology) overall = 0.1848
p_a_and_b = p_tech_and_west
p_b = p_west
p_a_given_b = p_a_and_b / p_b
print(f'P(Technology | West) via definition: {p_a_given_b:.4f}')
print(f'Matches direct calculation: {abs(p_a_given_b - p_tech_given_west) < 1e-9}')
P(Technology | West) via definition: 0.1870
Matches direct calculation: True
region_category = pd.crosstab(df['Region'], df['Category'], normalize='index')
print(region_category.round(4))
Category  Furniture  Office Supplies  Technology
Region                                          
Central      0.2071           0.6121      0.1808
East         0.2110           0.6011      0.1879
South        0.2049           0.6142      0.1809
West         0.2207           0.5923      0.1870

Part 4: Random Variables and Simulating with NumPy#

rng = np.random.default_rng(seed=42)
coin_flips = rng.integers(0, 2, size=10000)
p_heads = coin_flips.mean()
print(f'Simulated P(heads) over 10,000 flips: {p_heads:.4f}')
Simulated P(heads) over 10,000 flips: 0.4939
sampled_orders = rng.choice(df['Category'], size=5000, replace=True)
sampled_probs = pd.Series(sampled_orders).value_counts(normalize=True)
print('Simulated by resampling from the real category distribution:')
print(sampled_probs)
print()
print('Original real distribution:')
print(category_probs)
Simulated by resampling from the real category distribution:
Office Supplies    0.5998
Furniture          0.2110
Technology         0.1892
Name: proportion, dtype: float64

Original real distribution:
Category
Office Supplies    0.602962
Furniture          0.212227
Technology         0.184811
Name: proportion, dtype: float64

This resampling technique, drawing new simulated samples from real observed data, is called the bootstrap. It comes back in a much bigger way in video six, when we build confidence intervals without relying on any theoretical formula.

Part 5: The Law of Large Numbers#

import matplotlib.pyplot as plt

sample_sizes = np.arange(10, 5001, 10)
running_flips = rng.integers(0, 2, size=5000)
running_means = np.cumsum(running_flips) / np.arange(1, 5001)
plt.figure(figsize=(9, 4))
plt.plot(running_means, color='steelblue', linewidth=1)
plt.axhline(0.5, color='darkred', linestyle='--', label='True probability (0.5)')
plt.title('Simulated Coin Flip Probability Converging with Sample Size')
plt.xlabel('Number of Flips')
plt.ylabel('Running Probability of Heads')
plt.legend()
plt.show()
No description has been provided for this image
for n in [10, 100, 1000, 5000]:
    est = running_means[n - 1]
    print(f'After {n} flips: estimated P(heads) = {est:.4f}, error = {abs(est - 0.5):.4f}')
After 10 flips: estimated P(heads) = 0.6000, error = 0.1000
After 100 flips: estimated P(heads) = 0.5100, error = 0.0100
After 1000 flips: estimated P(heads) = 0.5120, error = 0.0120
After 5000 flips: estimated P(heads) = 0.4988, error = 0.0012

Wrap-Up: What You Learned#

  • Empirical probability: computing real probabilities directly from observed data.
  • The complement, addition, and multiplication-style rules, each verified two different ways on real data.
  • Conditional probability, and how to check whether two real variables look independent.
  • Simulating random processes with NumPy, including resampling real data with the bootstrap technique.
  • The law of large numbers, watched directly as a simulated estimate converges to the truth.
  • Video two moves from these general rules to formal discrete distributions: the binomial and the Poisson, applied to real daily order counts. Subscribe so it lands automatically see you there.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.