Mathew K Analytics

Lesson 29 · Probability and Statistics in python

Master Accuracy, Precision & Recall for Evaluating Classification Models in Python

Welcome! In this lesson, we will explore how to measure how well a classification model does its job. We will see how to use accuracy, precision, and recall…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Evaluating Classification Models: Accuracy, Precision, Recall#

Welcome! In this lesson, we will explore how to measure how well a classification model does its job.

We will see how to use accuracy, precision, and recall three key ideas for classification problems.

Real-world data from the Titanic disaster will help bring these ideas to life.

If you are new, do not worry! We will go step by step from the basics to hands-on Python code.

Let us begin!

# Suppress warnings for a cleaner output
import warnings; warnings.filterwarnings("ignore")

# Main libraries for data, plots, and evaluation
import pandas as pd
import numpy as np
import seaborn as sns
import matplotlib.pyplot as plt
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score, precision_score, recall_score, confusion_matrix
# Data setup: Load and peek at the Titanic dataset
titanic = sns.load_dataset("titanic")
print("Titanic shape:", titanic.shape)
titanic.head()
Titanic shape: (891, 15)
survived pclass sex age sibsp parch fare embarked class who adult_male deck embark_town alive alone
0 0 3 male 22.0 1 0 7.2500 S Third man True NaN Southampton no False
1 1 1 female 38.0 1 0 71.2833 C First woman False C Cherbourg yes False
2 1 3 female 26.0 0 0 7.9250 S Third woman False NaN Southampton yes True
3 1 1 female 35.0 1 0 53.1000 S First woman False C Southampton yes False
4 0 3 male 35.0 0 0 8.0500 S Third man True NaN Southampton no True

Why Evaluate Classification Models?#

We build classification models to predict things.

For example: Will a passenger on the Titanic survive?

It is important to measure how good our predictions are.

That is what evaluation metrics like accuracy, precision, and recall are for.

Let us learn what each one means!

Basic Definitions#

  • Accuracy: The percent of predictions that are correct.
  • Precision: The percent of positive predictions that are truly positive.
  • Recall: The percent of actual positives that we caught.

Think of "positive" as "survived" for this dataset.

We will come back to these after making a classifier!

# Let us prepare our data: Drop rows with missing age or embark_town
titanic_clean = titanic.dropna(subset=["age", "embark_town"]).copy()
print("Rows after cleaning:", titanic_clean.shape[0])
titanic_clean.head(3)
Rows after cleaning: 712
survived pclass sex age sibsp parch fare embarked class who adult_male deck embark_town alive alone
0 0 3 male 22.0 1 0 7.2500 S Third man True NaN Southampton no False
1 1 1 female 38.0 1 0 71.2833 C First woman False C Cherbourg yes False
2 1 3 female 26.0 0 0 7.9250 S Third woman False NaN Southampton yes True
# Choose features and the target (who survived)
X = titanic_clean[["pclass", "sex", "age", "fare"]]
y = titanic_clean["survived"]
print("X shape:", X.shape)
X.head(2)
X shape: (712, 4)
pclass sex age fare
0 3 male 22.0 7.2500
1 1 female 38.0 71.2833
# Convert 'sex' to numeric: male=1, female=0
X["sex"] = X["sex"].map({"male":1, "female":0})
X.head(2)
pclass sex age fare
0 3 1 22.0 7.2500
1 1 0 38.0 71.2833
# Split data: 70% train, 30% test. Always use random_state for reproducibility.
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.3, random_state=42
)
print("Train size:", X_train.shape[0], "Test size:", X_test.shape[0])
Train size: 498 Test size: 214
# Train a logistic regression model
model = LogisticRegression(max_iter=1000)
model.fit(X_train, y_train)
print("Ready to predict!")
Ready to predict!
# Predict on the test data
y_pred = model.predict(X_test)
print("First 10 predicted survival values:", y_pred[:10])
First 10 predicted survival values: [1 1 0 1 0 1 0 1 0 1]
# Calculate accuracy: What percent of test labels did we get right?
accuracy = accuracy_score(y_test, y_pred)
print("Accuracy:", round(accuracy, 3))
Accuracy: 0.78
# What about precision and recall?
precision = precision_score(y_test, y_pred)
recall = recall_score(y_test, y_pred)
print("Precision:", round(precision, 3))
print("Recall:", round(recall, 3))
Precision: 0.808
Recall: 0.641

Understanding Confusion Matrix#

A confusion matrix is a simple table that shows all prediction outcomes:

  • True Positive: Model said survived, person survived.
  • True Negative: Model said did not survive, person did not survive.
  • False Positive: Model said survived, person did not survive.
  • False Negative: Model said did not survive, person survived.

Let us see what the confusion matrix looks like for our test data.

# Show the confusion matrix
cm = confusion_matrix(y_test, y_pred)
print("Confusion matrix:\n", cm)
sns.heatmap(cm, annot=True, fmt="d", cmap="Blues")
plt.xlabel("Predicted")
plt.ylabel("Actual")
plt.title("Confusion Matrix")
plt.show()
Confusion matrix:
 [[108  14]
 [ 33  59]]
No description has been provided for this image

What If We Change the Threshold?#

Models usually output a probability. By default, we say above 0.5 means 'yes'.

But raising or lowering that threshold can change precision and recall.

You might want to find as many positives as possible (high recall), or be super sure before calling positive (high precision).

Let us try changing the threshold!

# Tune the threshold: Use 0.3 instead of 0.5
y_prob = model.predict_proba(X_test)[:,1]
custom_thresh = 0.3
y_pred_custom = (y_prob > custom_thresh).astype(int)
print("First 10 predictions (threshold 0.3):", y_pred_custom[:10])
First 10 predictions (threshold 0.3): [1 1 1 1 0 1 1 1 1 1]
# Precision and recall at threshold 0.3
precision_c = precision_score(y_test, y_pred_custom)
recall_c = recall_score(y_test, y_pred_custom)
print("Precision (0.3):", round(precision_c, 3))
print("Recall (0.3):", round(recall_c, 3))
Precision (0.3): 0.634
Recall (0.3): 0.772

Best Practices for Evaluating Classifiers#

  • Look at more than just accuracy. Explore precision and recall or use the F1 score.
  • Use a confusion matrix to understand types of mistakes.
  • Try different thresholds to balance recall and precision for your real-world problem.
  • Always use a clean test set to report real performance.

No model is perfect. Different domains require different choices.

# Your turn! Try your own threshold
answer = input("Type a threshold between 0 and 1 (for example, 0.6): ")
try:
    user_thresh = float(answer)
    y_pred_user = (y_prob > user_thresh).astype(int)
    print("Precision:", round(precision_score(y_test, y_pred_user), 3))
    print("Recall:", round(recall_score(y_test, y_pred_user), 3))
except:
    print("Please enter a valid number.")
    
Precision: 0.904
Recall: 0.511

Challenge Exercise#

  1. Change the features used to predict survival. Can you get better precision or recall?
  2. Try another scikit-learn classifier, such as RandomForestClassifier. Compare results.

Extra: Plot precision and recall as the threshold moves from 0 to 1.

Recap: What Did We Learn?#

Today you learned how to evaluate a classification model in Python.

  • Accuracy tells you what percent you got right.
  • Precision explains how careful your positives were.
  • Recall shows how many real positives you found.

You used the Titanic data and learned how to control model behavior with the threshold.

Congratulations!

Thank You!#

Thanks for learning with us!

If this video helped, please give a like or leave a comment.

Remember: Keep practicing, and you will build strong data skills.

See you next time!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.