Mathew K Analytics

Lesson 30 · Probability and Statistics in python

Understanding ROC Curves and AUC for Evaluating Classification Models in Python

Today we will explore how to evaluate classification models using ROC curves and AUC scores. We will use simple examples and hands-on code. By the end, you…

⬇ Download notebookOpen in Colab ↗

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Beginner Guide: ROC Curves and AUC in Python#

Today we will explore how to evaluate classification models using ROC curves and AUC scores.

We will use simple examples and hands-on code.

By the end, you will know how to plot and interpret ROC curves and calculate AUC!

# Suppress warnings (always use at the top for beginners)
import warnings
warnings.filterwarnings("ignore")

# Basic imports
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sns
# Data setup: Let us load the Titanic dataset
df = sns.load_dataset("titanic")
print("Shape:", df.shape)
df.head()
Shape: (891, 15)
survived pclass sex age sibsp parch fare embarked class who adult_male deck embark_town alive alone
0 0 3 male 22.0 1 0 7.2500 S Third man True NaN Southampton no False
1 1 1 female 38.0 1 0 71.2833 C First woman False C Cherbourg yes False
2 1 3 female 26.0 0 0 7.9250 S Third woman False NaN Southampton yes True
3 1 1 female 35.0 1 0 53.1000 S First woman False C Southampton yes False
4 0 3 male 35.0 0 0 8.0500 S Third man True NaN Southampton no True

Why study ROC curves and AUC?

ROC curves and AUC help us measure how well a model tells classes apart.

They are important any time you want to classify something, like predicting disease, fraud detection, or who survived the Titanic.

# Prepare the Titanic data: select only useful columns and drop missing values
cols_to_use = ["survived", "age", "fare", "sex", "pclass"]
df_clean = df[cols_to_use].dropna()
df_clean["sex"] = df_clean["sex"].map({"male":0, "female":1})
print("Clean shape:", df_clean.shape)
df_clean.head()
Clean shape: (714, 5)
survived age fare sex pclass
0 0 22.0 7.2500 0 3
1 1 38.0 71.2833 1 1
2 1 26.0 7.9250 1 3
3 1 35.0 53.1000 1 1
4 0 35.0 8.0500 0 3
# Split into features (X) and target (y)
X = df_clean.drop("survived", axis=1)
y = df_clean["survived"]
print("Features shape:", X.shape)
print("Target shape:", y.shape)
Features shape: (714, 4)
Target shape: (714,)
# Split into training and test sets
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)
print("Training set size:", X_train.shape[0])
print("Test set size:", X_test.shape[0])
Training set size: 499
Test set size: 215
# Fit a logistic regression model
from sklearn.linear_model import LogisticRegression
model = LogisticRegression(max_iter=500)
model.fit(X_train, y_train)
print("Model trained!")
Model trained!
# Predict probabilities instead of classes
y_probs = model.predict_proba(X_test)[:,1]
print(y_probs[:10])
[0.1691511  0.50521045 0.81012635 0.94742815 0.04799236 0.43823248
 0.48536465 0.60226047 0.5441697  0.59537501]

What is an ROC curve?

An ROC curve shows how well a model separates the two classes as you change the cutoff for deciding between them.

It is a useful way to judge models for things like disease detection or credit approval.

# Calculate ROC curve points
from sklearn.metrics import roc_curve
fpr, tpr, thresholds = roc_curve(y_test, y_probs)
print("Thresholds:", thresholds[:5])
Thresholds: [       inf 0.96317302 0.94136865 0.92927434 0.87592265]
# Plot the ROC curve
plt.figure(figsize=(6,6))
plt.plot(fpr, tpr, label="ROC curve")
plt.plot([0, 1], [0, 1], linestyle="--", color="gray", label="Chance")
plt.xlabel("False Positive Rate")
plt.ylabel("True Positive Rate")
plt.title("ROC Curve: Titanic Survival")
plt.legend()
plt.show()
No description has been provided for this image
# Calculate the AUC score
from sklearn.metrics import roc_auc_score
auc = roc_auc_score(y_test, y_probs)
print("AUC score:", auc)
AUC score: 0.8178170144462279

Interpreting ROC and AUC

  • An ROC curve closer to the top left means the model is better.
  • A perfect AUC is 1.0; random guessing gives 0.5.
  • Use these tools to compare models and pick the best one for your problem.
# Practice: Try a decision tree and compare!
from sklearn.tree import DecisionTreeClassifier
tree = DecisionTreeClassifier(random_state=42)
tree.fit(X_train, y_train)
y_probs_tree = tree.predict_proba(X_test)[:,1]

auc_tree = roc_auc_score(y_test, y_probs_tree)
print("Decision Tree AUC:", auc_tree)
Decision Tree AUC: 0.7166934189406099
# What if the classes are very imbalanced?
print("Survived count:")
print(df_clean["survived"].value_counts())
Survived count:
survived
0    424
1    290
Name: count, dtype: int64
# Mini-project: Simulate predictions with random noise
np.random.seed(2)
random_scores = np.random.rand(y_test.shape[0])
auc_random = roc_auc_score(y_test, random_scores)
print("Random predictions AUC:", auc_random)
Random predictions AUC: 0.46807561976101303
# Extra tip: Precision-recall curves for rare outcomes
from sklearn.metrics import precision_recall_curve, average_precision_score
precision, recall, pr_thresholds = precision_recall_curve(y_test, y_probs)

plt.figure(figsize=(6,6))
plt.plot(recall, precision, label="Precision-Recall curve")
plt.xlabel("Recall")
plt.ylabel("Precision")
plt.title("Precision-Recall Curve: Titanic Survival")
plt.legend()
plt.show()
No description has been provided for this image
# Challenge: What happens with a different threshold?
threshold = 0.3
y_pred_thresh = (y_probs > threshold).astype(int)
from sklearn.metrics import confusion_matrix
cm = confusion_matrix(y_test, y_pred_thresh)
print("Confusion matrix at threshold 0.3:")
print(cm)
Confusion matrix at threshold 0.3:
[[90 36]
 [18 71]]

Common troubles and tips

  • Always use the test set for AUC, not training!
  • If your data has very few positives, check both ROC and precision-recall plots.
  • Compare at least two models.
  • Check what happens if you change the decision threshold.

Well done! Let us recap:

  • ROC curves show how well the model sorts classes.
  • AUC is a single number that measures model quality.
  • Try different models, plot results, and be sure to check what matters for your task.

Stay curious and practice plotting ROC and AUC on new datasets!

If you enjoyed this, like and subscribe for more step-by-step guides.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.