Lesson 35 · Probability and Statistics in python
Bayesian Estimation Explained: Using SciPy and PyMC for Practical Inference in Python
Welcome! In this hands-on tutorial, you will explore probability, statistics, and Bayesian thinking using Python. No experience needed. By the end, youll…
- CourseProbability and Statistics in python
- Lesson35 of 35
- Video15 min
- FormatJupyter notebook · 12 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbProbability & Bayesian Estimation in Python: Complete Beginner Lesson#
Welcome! In this hands-on tutorial, you will explore probability, statistics, and Bayesian thinking using Python. No experience needed.
By the end, youll understand concepts like probability, mean, distributions, and you will try beginner-friendly Bayesian estimation with real data.
Let us start our journey!
import warnings
from statsmodels.tools.sm_exceptions import ConvergenceWarning, ValueWarning
warnings.filterwarnings("ignore", category=ValueWarning)
warnings.filterwarnings("ignore", category=ConvergenceWarning)
warnings.filterwarnings("ignore", category=RuntimeWarning)
# Import basic libraries for the whole notebook
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sns
np.random.seed(42)
What is Probability?#
Probability simply means how likely something is to happen. It is always a number between 0 (impossible) and 1 (certain).
Let us try simulating a coin flip.
# Simulate a fair coin flip 10 times
np.random.seed(1)
outcomes = np.random.choice(["Heads", "Tails"], size=10)
print("Coin flips:", outcomes)
# Count how many times we got heads
heads = np.sum(outcomes == "Heads")
tails = np.sum(outcomes == "Tails")
print("Heads:", heads, ", Tails:", tails)
Dice Rolling: A Simple Discrete Probability#
Let us roll a six-sided die 12 times and see the results.
# Simulate rolling a die 12 times
np.random.seed(2)
die_rolls = np.random.randint(1, 7, size=12)
print("Die rolls:", die_rolls)
# Count frequency of each face
values, counts = np.unique(die_rolls, return_counts=True)
for val, c in zip(values, counts):
print(f"Number {val}: {c} times")
Descriptive Statistics: Mean and Standard Deviation#
Descriptive statistics help you summarize a list of numbers, like your die rolls.
# Calculate mean and standard deviation
mean_value = np.mean(die_rolls)
std_value = np.std(die_rolls)
print("Mean:", mean_value)
print("Standard deviation:", std_value)
Handling Missing Data#
Missing data happens often in real datasets. Let us see how to spot it and deal with it safely.
# Create data with a missing value
data = pd.Series([5, 3, np.nan, 7, 6, np.nan])
print("Data:")
print(data)
# Find missing entries
missing = data.isnull()
print("Missing values:")
print(missing)
# Fill missing data with the mean
filled = data.fillna(data.mean())
print("After filling missing values:")
print(filled)
Discrete Distributions: Bernoulli, Binomial, Poisson#
Discrete distributions deal with outcomes you can count: successes, arrivals, events.
from scipy.stats import bernoulli, binom, poisson
# Bernoulli: Single coin flip (success/failure)
probs = bernoulli.pmf([0, 1], p=0.5)
print("Bernoulli PMF (p=0.5):", probs)
# Binomial: 10 trials, p=0.5, probability of 6 successes
binom_p = binom.pmf(6, n=10, p=0.5)
print("Binomial (n=10, p=0.5) probability of 6 successes:", binom_p)
# Poisson: Probability of 3 arrivals, average rate 2
poisson_p = poisson.pmf(3, mu=2)
print("Poisson (mu=2) probability of 3 events:", poisson_p)
Continuous Distributions: Normal, Exponential#
Continuous distributions model things that can take on any value in a range, like heights and waiting times.
from scipy.stats import norm, expon
# Normal distribution (bell curve)
x = np.linspace(-3, 3, 100)
y = norm.pdf(x, loc=0, scale=1)
plt.figure(figsize=(6, 2))
plt.plot(x, y, label="Normal PDF")
plt.title("Normal Distribution")
plt.legend()
plt.show()
# Exponential: Models waiting times
x2 = np.linspace(0, 6, 100)
y2 = expon.pdf(x2, scale=2)
plt.figure(figsize=(6, 2))
plt.plot(x2, y2, color="orange", label="Exponential PDF")
plt.title("Exponential Distribution")
plt.legend()
plt.show()
# Data setup: Load Titanic dataset, use seaborn version or fallback dataset if not available
try:
titanic = sns.load_dataset('titanic')
print('Shape:', titanic.shape)
print(titanic.head())
except Exception:
# Fallback: Use penguins if Titanic is not available
titanic = sns.load_dataset('penguins')
print('Titanic dataset not found. Fallback to penguins dataset.')
print('Shape:', titanic.shape)
print(titanic.head())
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



