Mathew K Analytics

Lesson 10 · Scikit-learn deep dive

Scikit-learn Tutorial #10: Classification Metrics

Video ten of the eighteen-part series: judging a classifier honestly, beyond plain accuracy. Precision, recall, F1, confusionmatrix, classificationreport,…

⬇ Download notebookOpen in Colab ↗

📓 Full notebook

Download .ipynb

Scikit-learn Deep-Dive, Video 10: Classification Metrics#

  • Video ten of the eighteen-part series: judging a classifier honestly, beyond plain accuracy.
  • Precision, recall, F1, confusion_matrix, classification_report, and ROC-AUC.
  • Let's get into it.

Part 1: accuracy_score - Why It Can Mislead on Imbalanced Data#

import numpy as np
from sklearn.metrics import accuracy_score
y_true_imbalanced = np.array([0]*95 + [1]*5)
y_pred_always_zero = np.zeros(100, dtype=int)
print(accuracy_score(y_true_imbalanced, y_pred_always_zero))
0.95

Part 2: confusion_matrix - TP, FP, TN, FN#

from sklearn.metrics import confusion_matrix
cm = confusion_matrix(y_true_imbalanced, y_pred_always_zero)
print(cm)
tn, fp, fn, tp = cm.ravel()
print('TN:', tn, 'FP:', fp, 'FN:', fn, 'TP:', tp)
[[95  0]
 [ 5  0]]
TN: 95 FP: 0 FN: 5 TP: 0

Part 3: precision_score and recall_score - the Tradeoff#

from sklearn.metrics import precision_score, recall_score
y_true_demo = np.array([1, 0, 1, 1, 0, 1, 0, 0])
y_pred_demo = np.array([1, 0, 0, 1, 1, 1, 0, 0])
print(round(precision_score(y_true_demo, y_pred_demo), 3))
print(round(recall_score(y_true_demo, y_pred_demo), 3))
0.75
0.75

Part 4: f1_score - Balancing Precision and Recall#

from sklearn.metrics import f1_score
print(round(f1_score(y_true_demo, y_pred_demo), 3))
high_precision_low_recall = np.array([0, 0, 0, 1, 0, 0, 0, 0])
print(round(precision_score(y_true_demo, high_precision_low_recall), 3))
print(round(recall_score(y_true_demo, high_precision_low_recall), 3))
print(round(f1_score(y_true_demo, high_precision_low_recall), 3))
0.75
1.0
0.25
0.4

Part 5: classification_report - All Metrics at Once#

from sklearn.metrics import classification_report
print(classification_report(y_true_demo, y_pred_demo, target_names=['negative', 'positive']))
              precision    recall  f1-score   support

    negative       0.75      0.75      0.75         4
    positive       0.75      0.75      0.75         4

    accuracy                           0.75         8
   macro avg       0.75      0.75      0.75         8
weighted avg       0.75      0.75      0.75         8

Part 6: predict_proba and decision_function#

from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LogisticRegression
Xc, yc = load_breast_cancer(return_X_y=True)
Xc_train, Xc_test, yc_train, yc_test = train_test_split(Xc, yc, test_size=0.25, random_state=42, stratify=yc)
clf = LogisticRegression(max_iter=5000).fit(Xc_train, yc_train)
probabilities = clf.predict_proba(Xc_test)
print(probabilities[:3].round(3))
[[0.019 0.981]
 [0.998 0.002]
 [0.177 0.823]]

Part 7: roc_curve and roc_auc_score#

from sklearn.metrics import roc_curve, roc_auc_score
positive_probs = probabilities[:, 1]
fpr, tpr, thresholds = roc_curve(yc_test, positive_probs)
auc = roc_auc_score(yc_test, positive_probs)
print(len(thresholds))
print(round(auc, 3))
12
0.996

Part 8: precision_recall_curve#

from sklearn.metrics import precision_recall_curve, average_precision_score
precisions, recalls, pr_thresholds = precision_recall_curve(yc_test, positive_probs)
avg_precision = average_precision_score(yc_test, positive_probs)
print(len(pr_thresholds))
print(round(avg_precision, 3))
143
0.998

Part 9: Multiclass Metrics - the average Parameter#

from sklearn.datasets import load_wine
Xw, yw = load_wine(return_X_y=True)
Xw_train, Xw_test, yw_train, yw_test = train_test_split(Xw, yw, test_size=0.25, random_state=42, stratify=yw)
wine_clf = LogisticRegression(max_iter=5000).fit(Xw_train, yw_train)
wine_preds = wine_clf.predict(Xw_test)
print(round(f1_score(yw_test, wine_preds, average='macro'), 3))
print(round(f1_score(yw_test, wine_preds, average='weighted'), 3))
print(round(f1_score(yw_test, wine_preds, average='micro'), 3))
0.952
0.955
0.956
c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\sklearn\linear_model\_logistic.py:473: ConvergenceWarning: lbfgs failed to converge after 5000 iteration(s) (status=1):
STOP: TOTAL NO. OF ITERATIONS REACHED LIMIT

Increase the number of iterations to improve the convergence (max_iter=5000).
You might also want to scale the data as shown in:
    https://scikit-learn.org/stable/modules/preprocessing.html
Please also refer to the documentation for alternative solver options:
    https://scikit-learn.org/stable/modules/linear_model.html#logistic-regression
  n_iter_i = _check_optimize_result(

Part 10: A Real Pattern - a Reusable evaluate_classifier Function#

def evaluate_classifier(y_true, y_pred):
    return {
        'accuracy': round(accuracy_score(y_true, y_pred), 3),
        'precision_macro': round(precision_score(y_true, y_pred, average='macro'), 3),
        'recall_macro': round(recall_score(y_true, y_pred, average='macro'), 3),
        'f1_macro': round(f1_score(y_true, y_pred, average='macro'), 3)
    }
print(evaluate_classifier(yw_test, wine_preds))
{'accuracy': 0.956, 'precision_macro': 0.967, 'recall_macro': 0.944, 'f1_macro': 0.952}

Wrap-Up: What You Learned#

  • accuracy_score can mislead badly on imbalanced data; a model predicting only the majority class can still score high.
  • confusion_matrix breaks predictions into true positives, false positives, true negatives, and false negatives.
  • precision measures how trustworthy a positive prediction is; recall measures how many actual positives were caught.
  • f1_score is the harmonic mean of precision and recall, penalizing lopsided performance.
  • classification_report prints precision, recall, F1, and support for every class at once.
  • predict_proba and decision_function give scores instead of hard labels, needed to build ROC and PR curves.
  • roc_curve and roc_auc_score trace and summarize true-positive rate versus false-positive rate across thresholds.
  • precision_recall_curve is often more informative than ROC on heavily imbalanced data.
  • macro, weighted, and micro averaging control how multiclass precision, recall, and F1 combine across classes.
  • That wraps up classification metrics. Next up: Regression Metrics - MSE, RMSE, MAE, and R-squared.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.