Lesson 30 · Probability and Statistics in python
Understanding ROC Curves and AUC for Evaluating Classification Models in Python
Today we will explore how to evaluate classification models using ROC curves and AUC scores. We will use simple examples and hands-on code. By the end, you…
- CourseProbability and Statistics in python
- Lesson30 of 35
- Video12 min
- FormatJupyter notebook · 15 code cells
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbBeginner Guide: ROC Curves and AUC in Python#
Today we will explore how to evaluate classification models using ROC curves and AUC scores.
We will use simple examples and hands-on code.
By the end, you will know how to plot and interpret ROC curves and calculate AUC!
# Suppress warnings (always use at the top for beginners)
import warnings
warnings.filterwarnings("ignore")
# Basic imports
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sns
# Data setup: Let us load the Titanic dataset
df = sns.load_dataset("titanic")
print("Shape:", df.shape)
df.head()
Why study ROC curves and AUC?
ROC curves and AUC help us measure how well a model tells classes apart.
They are important any time you want to classify something, like predicting disease, fraud detection, or who survived the Titanic.
# Prepare the Titanic data: select only useful columns and drop missing values
cols_to_use = ["survived", "age", "fare", "sex", "pclass"]
df_clean = df[cols_to_use].dropna()
df_clean["sex"] = df_clean["sex"].map({"male":0, "female":1})
print("Clean shape:", df_clean.shape)
df_clean.head()
# Split into features (X) and target (y)
X = df_clean.drop("survived", axis=1)
y = df_clean["survived"]
print("Features shape:", X.shape)
print("Target shape:", y.shape)
# Split into training and test sets
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)
print("Training set size:", X_train.shape[0])
print("Test set size:", X_test.shape[0])
# Fit a logistic regression model
from sklearn.linear_model import LogisticRegression
model = LogisticRegression(max_iter=500)
model.fit(X_train, y_train)
print("Model trained!")
# Predict probabilities instead of classes
y_probs = model.predict_proba(X_test)[:,1]
print(y_probs[:10])
What is an ROC curve?
An ROC curve shows how well a model separates the two classes as you change the cutoff for deciding between them.
It is a useful way to judge models for things like disease detection or credit approval.
# Calculate ROC curve points
from sklearn.metrics import roc_curve
fpr, tpr, thresholds = roc_curve(y_test, y_probs)
print("Thresholds:", thresholds[:5])
# Plot the ROC curve
plt.figure(figsize=(6,6))
plt.plot(fpr, tpr, label="ROC curve")
plt.plot([0, 1], [0, 1], linestyle="--", color="gray", label="Chance")
plt.xlabel("False Positive Rate")
plt.ylabel("True Positive Rate")
plt.title("ROC Curve: Titanic Survival")
plt.legend()
plt.show()
# Calculate the AUC score
from sklearn.metrics import roc_auc_score
auc = roc_auc_score(y_test, y_probs)
print("AUC score:", auc)
Interpreting ROC and AUC
- An ROC curve closer to the top left means the model is better.
- A perfect AUC is 1.0; random guessing gives 0.5.
- Use these tools to compare models and pick the best one for your problem.
# Practice: Try a decision tree and compare!
from sklearn.tree import DecisionTreeClassifier
tree = DecisionTreeClassifier(random_state=42)
tree.fit(X_train, y_train)
y_probs_tree = tree.predict_proba(X_test)[:,1]
auc_tree = roc_auc_score(y_test, y_probs_tree)
print("Decision Tree AUC:", auc_tree)
# What if the classes are very imbalanced?
print("Survived count:")
print(df_clean["survived"].value_counts())
# Mini-project: Simulate predictions with random noise
np.random.seed(2)
random_scores = np.random.rand(y_test.shape[0])
auc_random = roc_auc_score(y_test, random_scores)
print("Random predictions AUC:", auc_random)
# Extra tip: Precision-recall curves for rare outcomes
from sklearn.metrics import precision_recall_curve, average_precision_score
precision, recall, pr_thresholds = precision_recall_curve(y_test, y_probs)
plt.figure(figsize=(6,6))
plt.plot(recall, precision, label="Precision-Recall curve")
plt.xlabel("Recall")
plt.ylabel("Precision")
plt.title("Precision-Recall Curve: Titanic Survival")
plt.legend()
plt.show()
# Challenge: What happens with a different threshold?
threshold = 0.3
y_pred_thresh = (y_probs > threshold).astype(int)
from sklearn.metrics import confusion_matrix
cm = confusion_matrix(y_test, y_pred_thresh)
print("Confusion matrix at threshold 0.3:")
print(cm)
Common troubles and tips
- Always use the test set for AUC, not training!
- If your data has very few positives, check both ROC and precision-recall plots.
- Compare at least two models.
- Check what happens if you change the decision threshold.
Well done! Let us recap:
- ROC curves show how well the model sorts classes.
- AUC is a single number that measures model quality.
- Try different models, plot results, and be sure to check what matters for your task.
Stay curious and practice plotting ROC and AUC on new datasets!
If you enjoyed this, like and subscribe for more step-by-step guides.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



