Lesson 28 · Probability and Statistics in python
Logistic Regression Explained: Master Coefficients & Odds Ratios Easily
Welcome to today''s lesson! We will learn to read and interpret logistic regression outputs, focusing on coefficients and odds ratios. You will see hands-on…
- CourseProbability and Statistics in python
- Lesson28 of 35
- Video12 min
- FormatJupyter notebook · 14 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbInterpreting Coefficients and Odds Ratios in Logistic Regression#
Welcome to today''s lesson! We will learn to read and interpret logistic regression outputs, focusing on coefficients and odds ratios.
You will see hands-on examples using real data. Let''s begin!
import warnings
warnings.filterwarnings("ignore")
import pandas as pd
import numpy as np
import seaborn as sns
import matplotlib.pyplot as plt
# Data setup
titanic = sns.load_dataset("titanic")
print("Shape:", titanic.shape)
titanic.head()
What is Logistic Regression?#
Logistic regression is a method used to predict an outcome with two categories, like "survived" versus "did not survive".
Unlike linear regression, it predicts the chance (probability) that something happens.
# Let''s check missing values in the main columns we will use
print(titanic[["survived", "sex", "age", "pclass"]].isnull().sum())
# For simplicity, let''s drop rows with missing age
titanic_clean = titanic.dropna(subset=["age"]).copy()
print("Rows after dropping missing ages:", len(titanic_clean))
# Convert categorical variables to numeric codes
titanic_clean["sex_code"] = titanic_clean["sex"].map({"male": 0, "female": 1})
titanic_clean["pclass_code"] = titanic_clean["pclass"]
Odds, Probability, and Log-Odds#
Probability means the chance that something will happen. For example, 0.7 means 70 percent.
Odds are another way to describe chances: odds = p / (1-p). Log-odds are just the logarithm of the odds.
# Let''s see how probability and odds work
p = 0.8
odds = p / (1 - p)
log_odds = np.log(odds)
print(f"If probability is {p}, odds are {odds:.2f}, and log-odds are {log_odds:.2f}")
Setting Up Our Logistic Model#
We will predict survival on the Titanic using sex, class, and age.
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
# Set up X and y for modeling
X = titanic_clean[["sex_code", "pclass_code", "age"]]
y = titanic_clean["survived"]
# Split into train and test (80/20)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
# Fit the logistic regression model
lr = LogisticRegression(solver="lbfgs")
lr.fit(X_train, y_train)
# Look at model coefficients
features = ["sex_code", "pclass_code", "age"]
for name, coef in zip(features, lr.coef_[0]):
print(f"{name:10s}: {coef:+.3f}")
What Does a Coefficient Mean?#
A coefficient is how much the log-odds of survival changes with a one unit step in a feature, keeping other things the same.
For example, if "sex_code" is positive, being female helps survival compared to being male.
# Calculate odds ratios by exponentiating coefficients
odds_ratios = np.exp(lr.coef_[0])
for name, oratio in zip(features, odds_ratios):
print(f"{name:10s}: {oratio:.2f}")
Odds Ratio Interpretation#
An odds ratio above 1 means higher odds, below 1 means lower odds of survival.
For example, if "sex_code" odds ratio is 4.5, that means females are 4.5 times more likely to survive compared to males, all else equal.
# Predict survival probabilities for a sample passenger
example = pd.DataFrame({"sex_code": [0, 1], "pclass_code": [3, 1], "age": [25, 25]})
probs = lr.predict_proba(example)[:, 1]
print(f"Male, 3rd class, age 25 survives: {probs[0]:.2f}")
print(f"Female, 1st class, age 25 survives: {probs[1]:.2f}")
# See how odds ratio changes with age
age_diff = 10
ex1 = pd.DataFrame({"sex_code": [0], "pclass_code": [2], "age": [20]})
ex2 = pd.DataFrame({"sex_code": [0], "pclass_code": [2], "age": [20 + age_diff]})
prob1 = lr.predict_proba(ex1)[:, 1][0]
prob2 = lr.predict_proba(ex2)[:, 1][0]
change = prob2 / prob1 if prob1 > 0 else np.nan
print(f"A 10-year increase in age changes survival chance from {prob1:.2f} to {prob2:.2f}. Ratio: {change:.2f}")
# What if we use all test data? Check predicted vs. actual
y_pred = lr.predict(X_test)
accuracy = np.mean(y_pred == y_test)
print(f"Model accuracy on held-out data: {accuracy:.2%}")
Limitations and Best Practices#
Model coefficients show relationships, but real data may have hidden patterns, bias, or missing details.
Always check assumptions, and do not rely on one result alone.
# Practice: Try your own inputs!
sex = input("Enter 0 for male, 1 for female: ")
pclass = input("Enter 1 for 1st, 2 for 2nd, or 3 for 3rd class: ")
age = input("Enter passenger age: ")
your_data = pd.DataFrame({
"sex_code": [int(sex)],
"pclass_code": [int(pclass)],
"age": [float(age)]
})
prob = lr.predict_proba(your_data)[:, 1][0]
print(f"Your survival chance: {100*prob:.1f}%")
Recap: What We Learned About Coefficients and Odds Ratios#
You can now:
- Explain logistic regression coefficients in plain language
- Calculate odds ratios and interpret them
- Use these ideas to answer real questions about survival, medicine, and more!
Thanks for learning with us!
Want More Practice?#
Try these:
- Change the model by adding "fare" or "sibsp" as a feature.
- Build a logistic model for a different dataset, like predicting tips based on size, time, and day.
- Explain in your own words what an odds ratio below 1 means.
See you again soon!
Keep Exploring and Subscribe!#
For more hands-on Python and statistics, hit subscribe, leave questions, or suggest future topics in the comments. Happy learning!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



