Mathew K Analytics

Lesson 21 · Data Mining

Understanding ROC Curves, AUC, and Confusion Matrix for Classification Model Evaluation

Welcome! This week, we dive into classification model evaluation using the Telecom Customer Churn dataset. You will explore ROC curves, AUC, and confusion…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Week 56: ROC Curves, AUC, and Confusion Matrix Analysis#

Welcome! This week, we dive into classification model evaluation using the Telecom Customer Churn dataset.

You will explore ROC curves, AUC, and confusion matricescore tools for checking how well models perform.

These concepts help businesses decide when to trust a model's predictionsand spot errors early.

Ready? Let us get started!

# Data setup (Telecom Customer Churn Dataset)
import warnings; warnings.filterwarnings("ignore")
import pandas as pd
import numpy as np
np.random.seed(42)
url = 'https://raw.githubusercontent.com/IBM/telco-customer-churn-on-icp4d/master/data/Telco-Customer-Churn.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
(7043, 21)
   customerID  gender  SeniorCitizen Partner Dependents  tenure PhoneService  \
0  7590-VHVEG  Female              0     Yes         No       1           No   
1  5575-GNVDE    Male              0      No         No      34          Yes   
2  3668-QPYBK    Male              0      No         No       2          Yes   

      MultipleLines InternetService OnlineSecurity  ... DeviceProtection  \
0  No phone service             DSL             No  ...               No   
1                No             DSL            Yes  ...              Yes   
2                No             DSL            Yes  ...               No   

  TechSupport StreamingTV StreamingMovies        Contract PaperlessBilling  \
0          No          No              No  Month-to-month              Yes   
1          No          No              No        One year               No   
2          No          No              No  Month-to-month              Yes   

      PaymentMethod MonthlyCharges  TotalCharges Churn  
0  Electronic check          29.85         29.85    No  
1      Mailed check          56.95        1889.5    No  
2      Mailed check          53.85        108.15   Yes  

[3 rows x 21 columns]

Step 1: Data Cleaning and Preprocessing#

Real-world data often has missing or strange values.

Before any evaluation, let us clean the data for better results.

# Look for missing values
print(df.isnull().sum().sort_values(ascending=False).head(5))
customerID       0
gender           0
SeniorCitizen    0
Partner          0
Dependents       0
dtype: int64
# Remove rows with any missing values
df = df.dropna()
print(df.shape)
(7043, 21)
# Convert TotalCharges column to numeric (some are spaces)
df['TotalCharges'] = pd.to_numeric(df['TotalCharges'], errors='coerce')
df = df.dropna(subset=['TotalCharges'])
print(df['TotalCharges'].dtype)
float64

Step 2: Exploratory Data Analysis#

Let us understand what churn looks like in this dataset.

Visualization helps spot patterns early.

import matplotlib.pyplot as plt
# Plot the churn distribution
df['Churn'].value_counts().plot(kind='bar', color=['lightgreen','salmon'])
plt.title('Customer Churn Distribution')
plt.ylabel('Number of Customers')
plt.show()
No description has been provided for this image

Step 3: Classification Model Setup#

We now build a simple classifier to predict churn.

We will use logistic regressiona popular starting point.

from sklearn.model_selection import train_test_split
# Select features and encode categorical variables
features = ['SeniorCitizen','tenure','MonthlyCharges','TotalCharges']
X = df[features]
# Encode target as 1 (Yes) and 0 (No)
y = df['Churn'].map({'Yes':1, 'No':0})
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, random_state=42)
print('X_train shape:', X_train.shape)
print('y_train mean:', y_train.mean())
X_train shape: (5274, 4)
y_train mean: 0.26753886992794845
from sklearn.linear_model import LogisticRegression
# Train logistic regression on our data
clf = LogisticRegression(max_iter=500, random_state=42)
clf.fit(X_train, y_train)
# Predict probabilities on the test set
probs = clf.predict_proba(X_test)[:,1]

Step 4: Confusion Matrix Analysis#

A confusion matrix shows model successes and mistakes.

It counts true positives, false positives, true negatives, and false negatives.

from sklearn.metrics import confusion_matrix, classification_report
# Predict churn labels with standard 0.5 threshold
y_pred = (probs >= 0.5).astype(int)
cm = confusion_matrix(y_test, y_pred)
print(cm)
print(classification_report(y_test, y_pred))
[[1172  128]
 [ 254  204]]
              precision    recall  f1-score   support

           0       0.82      0.90      0.86      1300
           1       0.61      0.45      0.52       458

    accuracy                           0.78      1758
   macro avg       0.72      0.67      0.69      1758
weighted avg       0.77      0.78      0.77      1758

Step 5: ROC Curve and AUC#

The ROC curve shows the tradeoff between true and false positive rates.

AUC stands for Area Under the Curve. It measures overall performance.

Higher AUC values are better.

from sklearn.metrics import roc_curve, roc_auc_score
# Compute ROC curve points
fpr, tpr, thresholds = roc_curve(y_test, probs)
# Calculate AUC
auc_score = roc_auc_score(y_test, probs)
# Plot ROC curve
plt.plot(fpr, tpr, label='AUC = %.2f' % auc_score)
plt.plot([0,1],[0,1],'k--')
plt.xlabel('False Positive Rate')
plt.ylabel('True Positive Rate')
plt.title('ROC Curve')
plt.legend()
plt.show()
print('AUC score:', auc_score)
No description has been provided for this image
AUC score: 0.7996918038293584

Step 6: Varying Thresholds & Model Sensitivity#

Changing the probability threshold changes errors.

Let us see what happens if we lower the threshold.

# Try a lower threshold: more sensitive, more positives predicted
y_pred_low = (probs >= 0.3).astype(int)
cm_low = confusion_matrix(y_test, y_pred_low)
print(cm_low)
[[972 328]
 [142 316]]
# Input your own threshold to see its impact
user_threshold = float(input("Try a threshold (01, e.g. 0.6): "))
y_pred_user = (probs >= user_threshold).astype(int)
cm_user = confusion_matrix(y_test, y_pred_user)
print('Confusion matrix for threshold %.2f:' % user_threshold)
print(cm_user)
Confusion matrix for threshold 0.65:
[[1264   36]
 [ 371   87]]

Step 7: Recap and Best Practices#

  • Always check for missing data before training.

  • Look at ROC curves and confusion matrices, not just overall accuracy.

  • Experiment with different thresholds based on business needs.

  • Use AUC to compare classifiers fairly.

# Challenge: Try with a new feature
extra_feature = 'MonthlyCharges'
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
X_extra = scaler.fit_transform(df[[extra_feature]])
clf2 = LogisticRegression(random_state=42)
clf2.fit(X_extra, y)
probs2 = clf2.predict_proba(X_extra)[:,1]
auc2 = roc_auc_score(y, probs2)
print(f'AUC with only {extra_feature}:', round(auc2,2))
AUC with only MonthlyCharges: 0.62

Extra Challenge#

Try varying your feature selection or test out tree-based methods like DecisionTreeClassifier.

How do your confusion matrix and AUC results change?

End of Lesson Recap#

You learned how to:

  • Clean and prepare data;
  • Split data for modeling;
  • Interpret confusion matrices;
  • Plot and understand ROC curves;
  • Use AUC as a summary metric.

Keep experimenting and practicing!

If this was helpful, please subscribe and comment below with your favorite lesson moment!

See you in the next video.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.