Mathew K Analytics

Lesson 28 · Probability and Statistics in python

Logistic Regression Explained: Master Coefficients & Odds Ratios Easily

Welcome to today''s lesson! We will learn to read and interpret logistic regression outputs, focusing on coefficients and odds ratios. You will see hands-on…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Interpreting Coefficients and Odds Ratios in Logistic Regression#

Welcome to today''s lesson! We will learn to read and interpret logistic regression outputs, focusing on coefficients and odds ratios.

You will see hands-on examples using real data. Let''s begin!

import warnings
warnings.filterwarnings("ignore")
import pandas as pd
import numpy as np
import seaborn as sns
import matplotlib.pyplot as plt
# Data setup
titanic = sns.load_dataset("titanic")
print("Shape:", titanic.shape)
titanic.head()
Shape: (891, 15)
survived pclass sex age sibsp parch fare embarked class who adult_male deck embark_town alive alone
0 0 3 male 22.0 1 0 7.2500 S Third man True NaN Southampton no False
1 1 1 female 38.0 1 0 71.2833 C First woman False C Cherbourg yes False
2 1 3 female 26.0 0 0 7.9250 S Third woman False NaN Southampton yes True
3 1 1 female 35.0 1 0 53.1000 S First woman False C Southampton yes False
4 0 3 male 35.0 0 0 8.0500 S Third man True NaN Southampton no True

What is Logistic Regression?#

Logistic regression is a method used to predict an outcome with two categories, like "survived" versus "did not survive".

Unlike linear regression, it predicts the chance (probability) that something happens.

# Let''s check missing values in the main columns we will use
print(titanic[["survived", "sex", "age", "pclass"]].isnull().sum())
survived      0
sex           0
age         177
pclass        0
dtype: int64
# For simplicity, let''s drop rows with missing age
titanic_clean = titanic.dropna(subset=["age"]).copy()
print("Rows after dropping missing ages:", len(titanic_clean))
Rows after dropping missing ages: 714
# Convert categorical variables to numeric codes
titanic_clean["sex_code"] = titanic_clean["sex"].map({"male": 0, "female": 1})
titanic_clean["pclass_code"] = titanic_clean["pclass"]

Odds, Probability, and Log-Odds#

Probability means the chance that something will happen. For example, 0.7 means 70 percent.

Odds are another way to describe chances: odds = p / (1-p). Log-odds are just the logarithm of the odds.

# Let''s see how probability and odds work
p = 0.8
odds = p / (1 - p)
log_odds = np.log(odds)
print(f"If probability is {p}, odds are {odds:.2f}, and log-odds are {log_odds:.2f}")
If probability is 0.8, odds are 4.00, and log-odds are 1.39

Setting Up Our Logistic Model#

We will predict survival on the Titanic using sex, class, and age.

from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split

# Set up X and y for modeling
X = titanic_clean[["sex_code", "pclass_code", "age"]]
y = titanic_clean["survived"]

# Split into train and test (80/20)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
# Fit the logistic regression model
lr = LogisticRegression(solver="lbfgs")
lr.fit(X_train, y_train)
LogisticRegression()
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
# Look at model coefficients
features = ["sex_code", "pclass_code", "age"]
for name, coef in zip(features, lr.coef_[0]):
    print(f"{name:10s}: {coef:+.3f}")
    
sex_code  : +2.532
pclass_code: -1.249
age       : -0.043

What Does a Coefficient Mean?#

A coefficient is how much the log-odds of survival changes with a one unit step in a feature, keeping other things the same.

For example, if "sex_code" is positive, being female helps survival compared to being male.

# Calculate odds ratios by exponentiating coefficients
odds_ratios = np.exp(lr.coef_[0])
for name, oratio in zip(features, odds_ratios):
    print(f"{name:10s}: {oratio:.2f}")
    
sex_code  : 12.58
pclass_code: 0.29
age       : 0.96

Odds Ratio Interpretation#

An odds ratio above 1 means higher odds, below 1 means lower odds of survival.

For example, if "sex_code" odds ratio is 4.5, that means females are 4.5 times more likely to survive compared to males, all else equal.

# Predict survival probabilities for a sample passenger
example = pd.DataFrame({"sex_code": [0, 1], "pclass_code": [3, 1], "age": [25, 25]})
probs = lr.predict_proba(example)[:, 1]
print(f"Male, 3rd class, age 25 survives: {probs[0]:.2f}")
print(f"Female, 1st class, age 25 survives: {probs[1]:.2f}")
Male, 3rd class, age 25 survives: 0.10
Female, 1st class, age 25 survives: 0.95
# See how odds ratio changes with age
age_diff = 10
ex1 = pd.DataFrame({"sex_code": [0], "pclass_code": [2], "age": [20]})
ex2 = pd.DataFrame({"sex_code": [0], "pclass_code": [2], "age": [20 + age_diff]})
prob1 = lr.predict_proba(ex1)[:, 1][0]
prob2 = lr.predict_proba(ex2)[:, 1][0]
change = prob2 / prob1 if prob1 > 0 else np.nan
print(f"A 10-year increase in age changes survival chance from {prob1:.2f} to {prob2:.2f}. Ratio: {change:.2f}")
A 10-year increase in age changes survival chance from 0.33 to 0.24. Ratio: 0.74
# What if we use all test data? Check predicted vs. actual
y_pred = lr.predict(X_test)
accuracy = np.mean(y_pred == y_test)
print(f"Model accuracy on held-out data: {accuracy:.2%}")
Model accuracy on held-out data: 74.83%

Limitations and Best Practices#

Model coefficients show relationships, but real data may have hidden patterns, bias, or missing details.

Always check assumptions, and do not rely on one result alone.

# Practice: Try your own inputs!
sex = input("Enter 0 for male, 1 for female: ")
pclass = input("Enter 1 for 1st, 2 for 2nd, or 3 for 3rd class: ")
age = input("Enter passenger age: ")
your_data = pd.DataFrame({
    "sex_code": [int(sex)],
    "pclass_code": [int(pclass)],
    "age": [float(age)]
})
prob = lr.predict_proba(your_data)[:, 1][0]
print(f"Your survival chance: {100*prob:.1f}%")
Your survival chance: 83.3%

Recap: What We Learned About Coefficients and Odds Ratios#

You can now:

  • Explain logistic regression coefficients in plain language
  • Calculate odds ratios and interpret them
  • Use these ideas to answer real questions about survival, medicine, and more!

Thanks for learning with us!

Want More Practice?#

Try these:

  • Change the model by adding "fare" or "sibsp" as a feature.
  • Build a logistic model for a different dataset, like predicting tips based on size, time, and day.
  • Explain in your own words what an odds ratio below 1 means.

See you again soon!

Keep Exploring and Subscribe!#

For more hands-on Python and statistics, hit subscribe, leave questions, or suggest future topics in the comments. Happy learning!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.