Mathew K Analytics

Lesson 3 · Statistics for data analysts

Continuous Distributions & the Normal Curve | Statistics #3

Video three of the 15-part series: from counting outcomes to measuring them, the normal, uniform, and exponential distributions. Real mtcars fuel economy…

What you'll learn

Datasets used in this lesson

Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.

📓 Full notebook

Download .ipynb

Statistics for Data Analysts, Video 3: Continuous Distributions#

  • Video three of the 15-part series: from counting outcomes to measuring them, the normal, uniform, and exponential distributions.
  • Real mtcars fuel economy data, real AAPL daily returns, and real Superstore order gaps, plus one clearly-labeled simulated example.
  • Let's get into it.

Before You Start#

  • Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
  • You'll need pandas, NumPy, SciPy, and Matplotlib, all already installed if you've followed along.
  • Place mtcars.csv, aapl_ohlcv.csv, and superstore_sales.csv in the same folder as this notebook.
import pandas as pd
import numpy as np
from scipy import stats
import matplotlib.pyplot as plt

mtcars = pd.read_csv('mtcars.csv')
print(mtcars[['model', 'mpg']].head(3))
           model   mpg
0      Mazda RX4  21.0
1  Mazda RX4 Wag  21.0
2     Datsun 710  22.8

Part 1: The Normal Distribution#

mpg = mtcars['mpg']
mu, sigma = stats.norm.fit(mpg)
print(f'Real mtcars mpg: n={len(mpg)}')
print(f'Fitted normal mean (mu): {mu:.3f}')
print(f'Fitted normal std dev (sigma): {sigma:.3f}')
Real mtcars mpg: n=32
Fitted normal mean (mu): 20.091
Fitted normal std dev (sigma): 5.932
x = np.linspace(mpg.min() - 2, mpg.max() + 2, 200)
pdf = stats.norm.pdf(x, mu, sigma)
plt.figure(figsize=(9, 4))
plt.hist(mpg, bins=10, density=True, color='steelblue', edgecolor='white', alpha=0.7, label='Real mpg histogram')
plt.plot(x, pdf, color='darkred', linewidth=2, label='Fitted normal PDF')
plt.title('Real mtcars MPG vs. a Fitted Normal Curve')
plt.xlabel('Miles per Gallon')
plt.legend()
plt.show()
No description has been provided for this image
within_1sd = mpg.between(mu - sigma, mu + sigma).mean()
within_2sd = mpg.between(mu - 2 * sigma, mu + 2 * sigma).mean()
print(f'Real share within 1 std dev of the mean: {within_1sd:.1%} (rule expects ~68%)')
print(f'Real share within 2 std devs of the mean: {within_2sd:.1%} (rule expects ~95%)')
Real share within 1 std dev of the mean: 75.0% (rule expects ~68%)
Real share within 2 std devs of the mean: 93.8% (rule expects ~95%)

Part 2: A Real-World Stress Test - AAPL Daily Returns#

aapl = pd.read_csv('aapl_ohlcv.csv').dropna(subset=['Close'])
daily_returns = aapl['Close'].pct_change(fill_method=None).dropna()
mu_r, sigma_r = stats.norm.fit(daily_returns)
print(f'Real trading days: {len(daily_returns)}')
print(f'Real mean daily return: {mu_r:.5f}')
print(f'Real std dev of daily return: {sigma_r:.5f}')
Real trading days: 249
Real mean daily return: 0.00181
Real std dev of daily return: 0.01519
skew_r = stats.skew(daily_returns)
kurt_r = stats.kurtosis(daily_returns)
print(f'Real skewness: {skew_r:.3f} (0 = perfectly symmetric)')
print(f'Real excess kurtosis: {kurt_r:.3f} (0 = normal-like tails)')
Real skewness: 0.064 (0 = perfectly symmetric)
Real excess kurtosis: 2.116 (0 = normal-like tails)
plt.figure(figsize=(6, 6))
stats.probplot(daily_returns, dist='norm', plot=plt)
plt.title('Q-Q Plot: Real AAPL Daily Returns vs. Normal')
plt.show()
No description has been provided for this image

Part 3: The Uniform Distribution#

rng = np.random.default_rng(seed=11)
simulated_uniform = rng.uniform(low=0, high=10, size=5000)
print(f'Simulated sample mean: {simulated_uniform.mean():.3f} (theoretical: {(0 + 10) / 2:.3f})')
print(f'Simulated sample std: {simulated_uniform.std():.3f} (theoretical: {((10 - 0) ** 2 / 12) ** 0.5:.3f})')
Simulated sample mean: 4.978 (theoretical: 5.000)
Simulated sample std: 2.853 (theoretical: 2.887)
plt.figure(figsize=(9, 4))
plt.hist(simulated_uniform, bins=30, density=True, color='steelblue', edgecolor='white', alpha=0.7, label='Simulated draws')
plt.axhline(1 / 10, color='darkred', linewidth=2, label='Theoretical uniform PDF (flat)')
plt.title('Simulated Uniform(0, 10): Draws vs. Theoretical Flat PDF')
plt.xlabel('Value')
plt.legend()
plt.show()
No description has been provided for this image

Part 4: The Exponential Distribution - Real Order Gaps#

superstore = pd.read_csv('superstore_sales.csv')
superstore['Order Date'] = pd.to_datetime(superstore['Order Date'])
daily_counts = superstore.groupby(superstore['Order Date'].dt.date).size().sort_index()
busy_threshold = daily_counts.quantile(0.75)
busy_dates = pd.to_datetime(pd.Series(daily_counts[daily_counts > busy_threshold].index)).sort_values()
gaps_days = busy_dates.diff().dt.days.dropna()
print(f'Real busy-day threshold (75th percentile): {busy_threshold:.1f} order lines')
print(f'Real number of busy days: {len(busy_dates)}')
print(f'Real gap between busy days: mean={gaps_days.mean():.2f}, min={gaps_days.min():.0f}, max={gaps_days.max():.0f}')
Real busy-day threshold (75th percentile): 12.0 order lines
Real number of busy days: 267
Real gap between busy days: mean=5.41, min=1, max=64
loc, scale = stats.expon.fit(gaps_days, floc=0)
print(f'Fitted exponential scale (mean gap): {scale:.3f}')
x2 = np.linspace(0, gaps_days.max(), 200)
plt.figure(figsize=(9, 4))
plt.hist(gaps_days, bins=20, density=True, color='steelblue', edgecolor='white', alpha=0.7, label='Real busy-day gaps')
plt.plot(x2, stats.expon.pdf(x2, loc, scale), color='darkred', linewidth=2, label='Fitted exponential PDF')
plt.title('Real Gaps Between Busy Order Days vs. a Fitted Exponential')
plt.xlabel('Days Between Busy Days')
plt.legend()
plt.show()
Fitted exponential scale (mean gap): 5.406
No description has been provided for this image
p_within_2_days = stats.expon.cdf(2, loc, scale)
print(f'Modeled P(next busy day within 2 days) = {p_within_2_days:.4f}')
real_within_2 = (gaps_days <= 2).mean()
print(f'Real observed P(gap <= 2 days) = {real_within_2:.4f}')
Modeled P(next busy day within 2 days) = 0.3092
Real observed P(gap <= 2 days) = 0.4774

Part 5: PDF vs. CDF, Tied Together#

fig, axes = plt.subplots(1, 2, figsize=(11, 4))
x3 = np.linspace(mpg.min() - 2, mpg.max() + 2, 200)
axes[0].plot(x3, stats.norm.pdf(x3, mu, sigma), color='steelblue')
axes[0].set_title('PDF: Fitted Normal for Real MPG')
axes[1].plot(x3, stats.norm.cdf(x3, mu, sigma), color='darkred')
axes[1].set_title('CDF: Fitted Normal for Real MPG')
plt.tight_layout()
plt.show()
No description has been provided for this image

Wrap-Up: What You Learned#

  • The normal distribution, fit to real mtcars mpg and checked against the empirical rule.
  • An honest stress test on real AAPL daily returns: roughly symmetric, but genuinely fatter-tailed than a true normal, confirmed with skewness, kurtosis, and a Q-Q plot.
  • The uniform distribution, in a clearly simulated example, plus its mean and variance formulas.
  • The exponential distribution, fit to real gaps between busy Superstore order days, connecting directly back to the Poisson process from video two.
  • PDF and CDF, the two equivalent ways to describe any continuous distribution.
  • Video four builds on all of this with the Central Limit Theorem: why sample means converge to a normal shape almost regardless of the underlying real distribution. Subscribe so it lands automatically see you there.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.