Lesson 41 · Probability and Statistics in python
Comprehensive Statistical Analysis: Step-by-Step Capstone Project Case Study
In this hands-on lesson, we will explore practical probability and statistics techniques in Python through a full case study. You will learn how to load…
- CourseProbability and Statistics in python
- Lesson41 of 35
- Video34 min
- FormatJupyter notebook · 19 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbCapstone: Comprehensive Statistical Analysis Case Study#
In this hands-on lesson, we will explore practical probability and statistics techniques in Python through a full case study.
You will learn how to load data, analyze patterns, and test hypotheses using real datasets.
Let us get started on your journey to becoming comfortable with statistical thinking and Python!
# Data setup: Load Titanic dataset for analysis
import warnings; warnings.filterwarnings('ignore')
import pandas as pd
import seaborn as sns
import numpy as np
np.random.seed(42)
df = sns.load_dataset('titanic')
print('Shape:', df.shape)
df.head()
Step 1: Understanding Variables and Probability#
Probability means how likely something is to happen. The Titanic dataset lets us ask questions like: What is the chance a random passenger survived?
Let us start by exploring simple probabilities.
# Probability: What fraction of passengers survived?
survived_prob = df['survived'].mean()
print('Probability a random passenger survived: {:.2f}'.format(survived_prob))
Step 2: Descriptive Statistics#
Descriptive statistics show us the main features of data in a summary way.
Let us look at age and fare. We will check the mean, median, and standard deviation. These summarize typical values and spread.
# Calculate basic statistics for age and fare
print('Age: mean={:.1f}, median={:.1f}, std={:.1f}'.format(df['age'].mean(), df['age'].median(), df['age'].std()))
print('Fare: mean={:.2f}, median={:.2f}, std={:.2f}'.format(df['fare'].mean(), df['fare'].median(), df['fare'].std()))
# Handling missing values safely
print('Missing ages:', df['age'].isnull().sum())
df = df.copy()
df['age_filled'] = df['age'].fillna(df['age'].median())
print('Nulls after fill:', df['age_filled'].isnull().sum())
Step 3: Discrete Probability Distributions#
Discrete means values you can count. Examples are number of siblings or survived versus not survived.
The binomial distribution models the number of successes in a series of yes or no trials.
Let us simulate flipping a coin like the survival question.
# Simulate 10 coin flips, probability 0.5, 1000 times
import numpy as np
coin_flips = np.random.binomial(n=10, p=0.5, size=1000)
print('Simulated head counts (mean):', coin_flips.mean())
# Simulate number of siblings aboard to show Poisson distribution
from scipy.stats import poisson
sibs = df['sibsp']
print('Mean siblings:', sibs.mean())
pmf = poisson.pmf([0,1,2], mu=sibs.mean())
print('P(0): {:.2f}, P(1): {:.2f}, P(2): {:.2f}'.format(*pmf))
Step 4: Continuous Distributions#
Not all data is countedsome is measured, like age or fare. Distributions such as the Normal and Exponential help us model this kind of data.
Let us see if age follows a normal (bell curve) shape.
# Plot a histogram of age and compare to normal curve
import matplotlib.pyplot as plt
from scipy.stats import norm
ages = df['age_filled']
plt.hist(ages, bins=20, density=True, alpha=0.6, color='g')
xmin, xmax = ages.min(), ages.max()
x = np.linspace(xmin, xmax, 100)
p = norm.pdf(x, ages.mean(), ages.std())
plt.plot(x, p, 'k', linewidth=2)
plt.title('Age Distribution with Normal Curve')
plt.show()
# Exponential example: time between events can be modeled this way
from scipy.stats import expon
wait_times = expon.rvs(scale=2, size=1000)
plt.hist(wait_times, bins=30, color='c', alpha=0.7)
plt.title('Exponential Distribution of Simulated Wait Times')
plt.show()
Step 5: Sampling and the Central Limit Theorem (CLT)#
We often take samples because we cannot see all the data.
The Central Limit Theorem says averages of samples will look normal, even if the original data is not.
Let us try this with sample means for fare.
# Simulate sampling fare 500 times, each sample size 30
sample_means = [df['fare'].dropna().sample(30, replace=True).mean() for _ in range(500)]
plt.hist(sample_means, bins=20, color='m', alpha=0.7)
plt.title('Sampling Distribution of Sample Means (Fare)')
plt.xlabel('Sample mean')
plt.show()
Step 6: Confidence Intervals#
A confidence interval estimates a range where the true value likely falls.
We will compute a 95 percent confidence interval for the mean age.
# Calculate 95 percent CI for the mean age
import scipy.stats as stats
ages = df['age_filled']
n = len(ages)
mean = ages.mean()
sem = ages.std(ddof=1) / n**0.5
ci = stats.t.interval(0.95, n-1, loc=mean, scale=sem)
print('95 percent CI for mean age: ({:.2f}, {:.2f})'.format(*ci))
Step 7: Hypothesis Testing#
Hypothesis tests help us decide if differences are real or just random.
Let us check: Did women have different survival rates than men on the Titanic?
We will do a two-sample t-test.
# Test: Did women survive at a different rate than men?
survived_female = df[df['sex']=='female']['survived']
survived_male = df[df['sex']=='male']['survived']
t_stat, p_value = stats.ttest_ind(survived_female, survived_male)
print('T-statistic:', t_stat)
print('P-value:', p_value)
Step 8: Correlation and Covariance#
Correlation measures how two variables move together. Covariance is similar but not scaled.
Let us see if age and fare are related.
# Compute correlation and covariance for age and fare
corr = df['age_filled'].corr(df['fare'])
cov = df['age_filled'].cov(df['fare'])
print('Correlation:', round(corr,2))
print('Covariance:', round(cov,2))
Step 9: Simple Regression#
Regression helps us predict one variable from another.
We will predict fare from age using a simple linear model.
Let us see if older people paid more or less.
# Fit and plot linear regression: fare ~ age
from sklearn.linear_model import LinearRegression
mask = df['fare'].notnull() & df['age_filled'].notnull()
X = df.loc[mask, ['age_filled']]
y = df.loc[mask, 'fare']
model = LinearRegression().fit(X, y)
plt.scatter(X, y, alpha=0.4)
plt.plot(X, model.predict(X), color='red')
plt.xlabel('Age')
plt.ylabel('Fare')
plt.title('Regression: Predicting Fare from Age')
plt.show()
print('Model slope:', model.coef_[0], 'Intercept:', model.intercept_)
Step 10: Logistic Regression (Classification)#
Logistic regression predicts categories like survived or not.
Let us use age and class to predict if a passenger survived.
# Logistic regression predicting survival from age and class
from sklearn.linear_model import LogisticRegression
df['pclass_filled'] = df['pclass'].fillna(df['pclass'].mode()[0])
X = df[['age_filled', 'pclass_filled']].dropna()
y = df.loc[X.index, 'survived']
clf = LogisticRegression(max_iter=200).fit(X, y)
score = clf.score(X, y)
print('Prediction accuracy:', round(score, 2))
# Try logistic prediction for a new passenger
input_age = float(input('Enter a passenger age: '))
input_pclass = int(input('Enter ticket class (1, 2, 3): '))
proba = clf.predict_proba([[input_age, input_pclass]])[0,1]
print('Predicted survival probability: {:.2f}'.format(proba))
Step 11: Time Series Analysis#
Let us briefly see how to work with data over time using passengers from an airline dataset.
We will plot monthly totals and notice any trends.
# Load and plot airline passengers over time
url = 'https://raw.githubusercontent.com/jbrownlee/Datasets/master/airline-passengers.csv'
air = pd.read_csv(url)
air['Month'] = pd.to_datetime(air['Month'])
plt.plot(air['Month'], air['Passengers'])
plt.xlabel('Month')
plt.ylabel('Passenger Count')
plt.title('Monthly Airline Passengers')
plt.show()
Step 12: Bayesian Approach#
Bayesian methods update probabilities based on new evidence.
Let us use a Beta distribution to estimate the chance a random person survived.
We update our belief as we see more data.
# Estimate survival probability using Bayesian Beta
from scipy.stats import beta
success = df['survived'].sum()
fail = df['survived'].size - success
x = np.linspace(0, 1, 200)
prior_alpha = prior_beta = 1
posterior = beta(success + prior_alpha, fail + prior_beta)
plt.plot(x, posterior.pdf(x))
plt.xlabel('Survival Probability')
plt.ylabel('Density')
plt.title('Bayesian Posterior for Survival Probability')
plt.show()
Step 13: Mini-Project - Simulation#
Let us model an emergency scenarioa lifeboat that can only save a fixed number of people.
How many would be left if picking at random? Let us simulate and see.
# Simulate saving 200 people from the Titanic at random, 1000 times
pop_size = len(df)
boat_size = 200
num_trials = 1000
left_behind = [pop_size - np.random.choice(pop_size, boat_size, replace=False).size for _ in range(num_trials)]
plt.hist(left_behind, bins=15, color='orange')
plt.xlabel('People Left Behind')
plt.ylabel('Frequency')
plt.title('Simulation: Titanic Rescue')
plt.show()
Step 14: Mini-Project - Analyze Real Data#
Let us do something practical: investigate if higher ticket price meant a better chance of survival.
You will group passengers by fare bands and check survival rates.
# Group passengers into fare bands and calculate survival by band
df['fare_band'] = pd.qcut(df['fare'], 4, labels=['Low', 'Mid-low', 'Mid-high', 'High'])
grouped = df.groupby('fare_band')['survived'].mean()
print('Survival rate by fare band:')
print(grouped)
Step 15: Best Practices and Troubleshooting#
When you work with data: always check for missing values, outliers, and suspicious zeros.
Keep your code tidy and document your thoughts as you work.
Exploratory analysis saves you from mistakes later!
Step 16: Extra Tips#
Try new datasets from seaborn, sklearn, or online for more practice.
Break big questions into small steps.
Read code out loud to catch mistakes and understand logic better.
Step 17: Challenge Exercises#
Try comparing survival by age groupwho had better odds: kids, adults, or seniors?
Rerun your logistic regression on new featureswhat changes?
Plot a time series of daily temperatures from another dataset.
Share your results with a friend or online community!
Recap: What You Learned in This Capstone#
You loaded real data, discovered hidden trends, tested ideas, and built models.
You can now turn raw numbers into knowledge.
Keep practicing and asking questions. Data science rewards curiosity!
Thank you for joining this lesson!#
If you enjoyed it, like the video and subscribe for more beginner-to-expert data science guides.
Which topic would you like to learn next? Leave a comment below!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



