Mathew K Analytics

Lesson 9 · Probability and Statistics in python

Master Bayes Theorem: Real-Life Examples & Step-by-Step Probability Explained

In this lesson, you will build real intuition for probability, statistics, and Bayes' theorem. We will use Python to play with coins, dice, and real-life…

⬇ Download notebookOpen in Colab ↗

📓 Full notebook

Download .ipynb

Welcome to Probability & Statistics in Python!#

In this lesson, you will build real intuition for probability, statistics, and Bayes' theorem.

We will use Python to play with coins, dice, and real-life data like the Titanic.

Let's get curious!

# Suppress warnings to keep outputs clean
import warnings
from statsmodels.tools.sm_exceptions import ConvergenceWarning, ValueWarning
warnings.filterwarnings("ignore", category=ValueWarning)
warnings.filterwarnings("ignore", category=ConvergenceWarning)
warnings.filterwarnings("ignore", category=RuntimeWarning)
import warnings as _w; _w.filterwarnings("ignore")  # Extra line for broad suppression

What is Probability?#

Probability is a way to describe how likely something is to happen, between 0 (never) and 1 (always).

Example: The chance of flipping a coin and getting heads is 0.5.

Let us try it out!

# Let's simulate flipping a coin 10 times
import numpy as np
coin_flips = np.random.choice(['Heads', 'Tails'], size=10)
print("Coin flips:", coin_flips)
Coin flips: ['Heads' 'Tails' 'Tails' 'Tails' 'Heads' 'Heads' 'Heads' 'Tails' 'Tails'
 'Tails']

Dice Rolls: Classic Probability#

Rolling dice is a great way to see probability in action.

Let's roll a six-sided die 20 times.

dice_rolls = np.random.choice(np.arange(1, 7), size=20)
print("Dice rolls:", dice_rolls)
Dice rolls: [3 3 2 4 5 2 6 3 1 2 1 1 1 3 6 3 5 1 2 2]

Descriptive Statistics: Mean, Median, and Mode#

Descriptive statistics summarize data in simple numbers.

  • Mean: The average. Add up and divide.
  • Median: The middle value.
  • Mode: The value that appears most often.

Let us calculate them!

import scipy.stats as stats
mean_roll = np.mean(dice_rolls)
median_roll = np.median(dice_rolls)
mode_result = stats.mode(dice_rolls, keepdims=True)
mode_roll = mode_result.mode[0] if hasattr(mode_result, "mode") else mode_result[0]
print("Mean:", mean_roll)
print("Median:", median_roll)
print("Mode:", mode_roll)
Mean: 2.8
Median: 2.5
Mode: 1

Working with Real Data: The Titanic#

Let us work with data from the Titanic.

This famous dataset records info about the survival of passengers.

Tip: Real datasets sometimes have missing values. We will handle those safely.

# Data setup
import pandas as pd
import seaborn as sns
titanic = sns.load_dataset("titanic")
print("Shape:", titanic.shape)
print(titanic.head())
Shape: (891, 15)
   survived  pclass     sex   age  sibsp  parch     fare embarked  class  \
0         0       3    male  22.0      1      0   7.2500        S  Third   
1         1       1  female  38.0      1      0  71.2833        C  First   
2         1       3  female  26.0      0      0   7.9250        S  Third   
3         1       1  female  35.0      1      0  53.1000        S  First   
4         0       3    male  35.0      0      0   8.0500        S  Third   

     who  adult_male deck  embark_town alive  alone  
0    man        True  NaN  Southampton    no  False  
1  woman       False    C    Cherbourg   yes  False  
2  woman       False  NaN  Southampton   yes   True  
3  woman       False    C  Southampton   yes  False  
4    man        True  NaN  Southampton    no   True  
# Check for missing values
print(titanic.isnull().sum())
survived         0
pclass           0
sex              0
age            177
sibsp            0
parch            0
fare             0
embarked         2
class            0
who              0
adult_male       0
deck           688
embark_town      2
alive            0
alone            0
dtype: int64
# Fill missing age values with the median age
titanic['age'] = titanic['age'].fillna(titanic['age'].median())
# Drop any rows missing 'embarked'
titanic = titanic.dropna(subset=['embarked'])
# Let us look at the chance of survival
survival_rate = titanic['survived'].mean()
print("Survival rate:", survival_rate)
Survival rate: 0.38245219347581555

Discrete Probability Distributions#

Some probabilities describe single events you can count, like number of heads in coins, or survivors.

  • Bernoulli: 0 or 1 outcomes, like survive/ not survive
  • Binomial: Many yes/no events, like flipping many coins
  • Poisson: Counts of rare events, like accidents

Let's see a binomial in action with coins!

# Simulating 1000 coin flips
flips = np.random.binomial(n=1, p=0.5, size=1000)
print("Number of heads:", np.sum(flips))
Number of heads: 491

Continuous Probability Distributions#

Some probabilities deal with measurements like height or waiting times.

  • Normal distribution: The bell curve, for heights, test scores, and more.
  • Exponential: For times between random events, like bus arrivals.

Let us build and plot some!

import matplotlib.pyplot as plt
normal_data = np.random.normal(loc=0, scale=1, size=500)
plt.hist(normal_data, bins=30, alpha=0.7, color='skyblue')
plt.title("Normal Distribution Example")
plt.show()
No description has been provided for this image
# Exponential distribution demo
exp_data = np.random.exponential(scale=2.0, size=500)
plt.hist(exp_data, bins=30, color='salmon', alpha=0.8)
plt.title("Exponential Distribution")
plt.show()
No description has been provided for this image

Sampling and the Central Limit Theorem#

If you take many random samples from data, the averages themselves form a bell curve. This fact is the Central Limit Theorem.

Let's try it by sampling from our Titanic ages.

# Take 100 samples of 20 ages each, find the mean each time
sample_means = [titanic['age'].sample(20, replace=True).mean() for _ in range(100)]
plt.hist(sample_means, bins=20, color='mediumpurple', alpha=0.8)
plt.title("Distribution of Sample Means (CLT Example)")
plt.xlabel("Sample Mean Age")
plt.show()
No description has been provided for this image

Confidence Intervals: Estimating with Uncertainty#

A confidence interval gives a range for an estimate, like a 'margin of error' around a survey result.

Let us build a 95% confidence interval for the average age on the Titanic.

from scipy.stats import sem, t
ages = titanic['age'].dropna()
mean_age = ages.mean()
ci = t.interval(0.95, len(ages)-1, loc=mean_age, scale=sem(ages))
print(f"95% confidence interval for mean age: {ci}")
95% confidence interval for mean age: (np.float64(28.460421273737143), np.float64(30.169882438298846))

Hypothesis Testing: Did Survival Depend on Sex?#

A hypothesis test checks if a pattern in data is likely real or just luck.

Let us test if survival rates differed for males and females on the Titanic.

survived_f = titanic[titanic['sex']=='female']['survived']
survived_m = titanic[titanic['sex']=='male']['survived']
from scipy.stats import ttest_ind
t_stat, p_value = ttest_ind(survived_f, survived_m, nan_policy="omit")
print(f"p-value: {p_value}")
p-value: 6.682012140613411e-69

Correlation and Covariance#

Correlation describes how two things move together. Scores close to 1 mean strong connection, 0 means no link.

Covariance measures if two variables go up and down together, but is harder to compare across datasets.

Let us look at age and fare on the Titanic.

corr = titanic['age'].corr(titanic['fare'])
covar = titanic['age'].cov(titanic['fare'])
print(f"Correlation: {corr}")
print(f"Covariance: {covar}")
Correlation: 0.09370714330015861
Covariance: 60.470974586240494

Simple Linear Regression: Predict Fare from Age#

Regression lets us predict one value from another.

Let us predict what fare someone paid based on their age.

from sklearn.linear_model import LinearRegression
X = titanic[['age']].values
y = titanic['fare'].values
reg = LinearRegression()
reg.fit(X, y)
pred_fare = reg.predict([[30]])
print(f"Predicted fare for age 30: {pred_fare[0]:.2f}")
Predicted fare for age 30: 32.34

Logistic Regression: Can We Predict Survival?#

Logistic regression predicts yes/no outcomes from features.

Let us try predicting who survived on the Titanic.

We will use age, sex, and class.

from sklearn.model_selection import train_test_split
from sklearn.linear_model import LogisticRegression
df = titanic.copy()
df['sex'] = (df['sex'] == 'male').astype(int)
X = df[['age', 'sex', 'pclass']].values
y = df['survived'].values
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
clf = LogisticRegression(max_iter=200)
clf.fit(X_train, y_train)
score = clf.score(X_test, y_test)
print(f"Test set accuracy: {score:.2f}")
Test set accuracy: 0.79

Time Series Basics: Airline Passengers Over Time#

A time series tracks values over regular intervals.

Let us plot the number of airline passengers each month, from a real CSV.

air_url = "https://raw.githubusercontent.com/jbrownlee/Datasets/master/airline-passengers.csv"
airline = pd.read_csv(air_url)
plt.plot(pd.to_datetime(airline['Month']), airline['Passengers'])
plt.title("Monthly Airline Passengers")
plt.xlabel("Month")
plt.ylabel("Passengers")
plt.show()
No description has been provided for this image

An Introduction to Bayesian Thinking#

Bayesian probability lets us update beliefs with evidence.

Bayes' theorem is a simple math rule for this:

P(A | B) = (P(B | A) * P(A)) / P(B)

  • P(A | B): Probability A is true given B is observed.

Let us see a simple example with spam emails.

# Let's work through Bayes' theorem: spam emails
# Suppose 20% of emails are spam (P(Spam)=0.2).
# 95% of spam emails contain 'win' (P(Win|Spam)=0.95),
# and 5% of non-spam emails contain 'win' (P(Win|NotSpam)=0.05).
# If you see 'win', what's the chance it's spam?
P_spam = 0.2
P_win_given_spam = 0.95
P_win_given_notspam = 0.05
P_notspam = 1 - P_spam
P_win = P_win_given_spam*P_spam + P_win_given_notspam*P_notspam
P_spam_given_win = (P_win_given_spam*P_spam) / P_win
print(f"P(Spam | Win): {P_spam_given_win:.2f}")
P(Spam | Win): 0.83

Mini-Project Part 1: Simulate a Probability Experiment#

Let us simulate a medical test, where most people are healthy.

  • 1% of people have the illness.
  • The test is 99% accurate if ill, 95% accurate if healthy.

We will simulate 10000 people and see the real odds after a positive test.

N = 10000
ill = np.random.choice([1, 0], size=N, p=[0.01, 0.99])
test_pos = np.zeros(N)
for i in range(N):
    if ill[i]:
        test_pos[i] = np.random.rand() < 0.99  # True positive
    else:
        test_pos[i] = np.random.rand() < 0.05  # False positive
# Bayesian answer: P(Ill | Test+) = 
num_positive = np.sum(test_pos)
num_ill_positive = np.sum((ill == 1) & (test_pos == 1))
prob = num_ill_positive / num_positive
print(f"Chance you are really ill if positive: {prob:.3f}")
Chance you are really ill if positive: 0.168

Mini-Project Part 2: Analyze a Real Dataset with Bayes' Rule#

Now let us estimate the probability a third-class Titanic passenger survived, using Bayes' rule.

We will compute:

  • P(Survived | Class=3)
num_class3 = (titanic['pclass']==3).sum()
num_surv_class3 = ((titanic['pclass']==3) & (titanic['survived']==1)).sum()
prob_surv_class3 = num_surv_class3 / num_class3
print(f"Chance a 3rd-class passenger survived: {prob_surv_class3:.2f}")
Chance a 3rd-class passenger survived: 0.24

Best Practices and Troubleshooting#

  • Always check for missing data
  • Use graphs to spot patterns and mistakes
  • Start simple, build up slowly
  • Remember: even experts go back and check their steps!

Do not panic if your output surprises you. That is part of discovery.

Extra Tips#

  • Explore with print() and .head()
  • Try using .groupby() in pandas
  • Use np.random.seed() for repeatable randomness
  • Read docs and search for example code if stuck

Keep playing and practicing!

Challenge Exercises#

  • Find the survival rate for children under 12
  • Test if fare predicts survival better than class
  • Simulate 500 dice rolls and find the mean, then compare to 1,000, 10,000 rolls
  • Make your own time series with random numbers and plot it!

Share your answers with a friend or online!

Recap: What You Learned#

  • Probability and statistics basics
  • How to explore and clean data
  • Classic and Bayesian ways to reason about chance
  • Real data analysis with code

Practice and experimenting will help you grow confident over time.

Thanks for Learning Together!#

If you enjoyed this, hit like and subscribe on YouTube to support more lessons.

Ask questions in the comments or share your projects below.

See you in the next data adventure!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.