Mathew K Analytics

Lesson 20 · Data Mining

Understanding Classifier Evaluation: Accuracy, Precision, Recall, and F1-Score Explained

Welcome to this lesson where we explore how to evaluate machine learning classifiers. We will talk about concepts, write code, and do hands-on exercises.…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Week 56: Evaluating ClassifiersAccuracy, Precision, Recall, and F1-Score#

Welcome to this lesson where we explore how to evaluate machine learning classifiers. We will talk about concepts, write code, and do hands-on exercises.

These metrics help us understand how well our models perform in real life.

We will use the Telecom Customer Churn dataset and the Titanic dataset. Let's start with a quick overview of why evaluation matters!

Why do we evaluate classifiers?#

Building a model is only half the job. Evaluation tells us if predictions are meaningful or misleading.

Wrong metrics can hide problems.

Today we will avoid common pitfalls!

# Suppress warnings for a smoother learning experience
import warnings; warnings.filterwarnings('ignore')
import numpy as np
np.random.seed(42)
# Data setup (Telecom Customer Churn Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/IBM/telco-customer-churn-on-icp4d/master/data/Telco-Customer-Churn.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
(7043, 21)
   customerID  gender  SeniorCitizen Partner Dependents  tenure PhoneService  \
0  7590-VHVEG  Female              0     Yes         No       1           No   
1  5575-GNVDE    Male              0      No         No      34          Yes   
2  3668-QPYBK    Male              0      No         No       2          Yes   

      MultipleLines InternetService OnlineSecurity  ... DeviceProtection  \
0  No phone service             DSL             No  ...               No   
1                No             DSL            Yes  ...              Yes   
2                No             DSL            Yes  ...               No   

  TechSupport StreamingTV StreamingMovies        Contract PaperlessBilling  \
0          No          No              No  Month-to-month              Yes   
1          No          No              No        One year               No   
2          No          No              No  Month-to-month              Yes   

      PaymentMethod MonthlyCharges  TotalCharges Churn  
0  Electronic check          29.85         29.85    No  
1      Mailed check          56.95        1889.5    No  
2      Mailed check          53.85        108.15   Yes  

[3 rows x 21 columns]
# Check for missing values
print(df.isnull().sum().sort_values(ascending=False).head(5))
customerID       0
gender           0
SeniorCitizen    0
Partner          0
Dependents       0
dtype: int64
# Basic cleanup: Drop customerID and rows with missing values
df = df.drop(['customerID'], axis=1)
df = df.dropna()
print(df.shape)
(7043, 20)
# Convert categorical columns to numbers
for c in df.select_dtypes('object').columns:
    df[c] = df[c].astype('category').cat.codes
print(df.head(3))
   gender  SeniorCitizen  Partner  Dependents  tenure  PhoneService  \
0       0              0        1           0       1             0   
1       1              0        0           0      34             1   
2       1              0        0           0       2             1   

   MultipleLines  InternetService  OnlineSecurity  OnlineBackup  \
0              1                0               0             2   
1              0                0               2             0   
2              0                0               2             2   

   DeviceProtection  TechSupport  StreamingTV  StreamingMovies  Contract  \
0                 0            0            0                0         0   
1                 2            0            0                0         1   
2                 0            0            0                0         0   

   PaperlessBilling  PaymentMethod  MonthlyCharges  TotalCharges  Churn  
0                 1              2           29.85          2505      0  
1                 0              3           56.95          1466      0  
2                 1              3           53.85           157      1  
# Split the data into features and labels
X = df.drop('Churn', axis=1)
y = df['Churn']
# Train-test split
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, random_state=42)
print(X_train.shape, X_test.shape)
(5282, 19) (1761, 19)
# Train a simple classifier: Logistic Regression
from sklearn.linear_model import LogisticRegression
model = LogisticRegression(max_iter=500)
model.fit(X_train, y_train)
LogisticRegression(max_iter=500)
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
# Make predictions on the test set
y_pred = model.predict(X_test)

Understanding Accuracy#

Accuracy is the fraction of correct predictions out of all predictions.

Works well when classes are balanced. Can be misleading for imbalanced data!

Let us compute it next.

# Calculate accuracy
from sklearn.metrics import accuracy_score
accuracy = accuracy_score(y_test, y_pred)
print('Accuracy:', accuracy)
Accuracy: 0.8126064735945485

Precision and Recall#

Precision answers: Of all predicted positives, how many were truly positive? Recall answers: Of all actual positives, how many did the model find?

Both matter if making mistakes is expensive!

# Calculate precision and recall
from sklearn.metrics import precision_score, recall_score
precision = precision_score(y_test, y_pred)
recall = recall_score(y_test, y_pred)
print('Precision:', precision)
print('Recall:', recall)
Precision: 0.6955380577427821
Recall: 0.5532359081419624

What is F1-Score?#

F1-Score combines precision and recall into one metric. It is their harmonic mean.

Best when you want a balance between both.

# Calculate F1-Score
from sklearn.metrics import f1_score
f1 = f1_score(y_test, y_pred)
print('F1-Score:', f1)
F1-Score: 0.6162790697674418
# Confusion matrix for deeper insight
from sklearn.metrics import confusion_matrix
cm = confusion_matrix(y_test, y_pred)
print(cm)
[[1166  116]
 [ 214  265]]
# Show classification report
from sklearn.metrics import classification_report
print(classification_report(y_test, y_pred))
              precision    recall  f1-score   support

           0       0.84      0.91      0.88      1282
           1       0.70      0.55      0.62       479

    accuracy                           0.81      1761
   macro avg       0.77      0.73      0.75      1761
weighted avg       0.80      0.81      0.81      1761

Switching datasets: Titanic#

Let us try a totally different dataset to check what happens when class sizes change.

Ready? Here we go!

# Data setup (Titanic Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
df2 = pd.read_csv(url)
print(df2.shape)
print(df2.head(3))
(891, 12)
   PassengerId  Survived  Pclass  \
0            1         0       3   
1            2         1       1   
2            3         1       3   

                                                Name     Sex   Age  SibSp  \
0                            Braund, Mr. Owen Harris    male  22.0      1   
1  Cumings, Mrs. John Bradley (Florence Briggs Th...  female  38.0      1   
2                             Heikkinen, Miss. Laina  female  26.0      0   

   Parch            Ticket     Fare Cabin Embarked  
0      0         A/5 21171   7.2500   NaN        S  
1      0          PC 17599  71.2833   C85        C  
2      0  STON/O2. 3101282   7.9250   NaN        S  
# Prepare basic Titanic features
df2 = df2[['Survived','Pclass','Sex','Age','SibSp','Fare']]
df2 = df2.dropna()
df2['Sex'] = df2['Sex'].astype('category').cat.codes
X2 = df2.drop('Survived', axis=1)
y2 = df2['Survived']
# Train-test split for Titanic
from sklearn.model_selection import train_test_split
X2_train, X2_test, y2_train, y2_test = train_test_split(X2, y2, test_size=0.2, random_state=42)
# Train and evaluate a logistic regression model for Titanic
from sklearn.linear_model import LogisticRegression
model2 = LogisticRegression(max_iter=300)
model2.fit(X2_train, y2_train)
y2_pred = model2.predict(X2_test)
from sklearn.metrics import classification_report
print(classification_report(y2_test, y2_pred))
              precision    recall  f1-score   support

           0       0.79      0.82      0.80        87
           1       0.70      0.66      0.68        56

    accuracy                           0.76       143
   macro avg       0.74      0.74      0.74       143
weighted avg       0.75      0.76      0.75       143

Practice and Try It Yourself!#

Switch back to your favorite dataset. Compute accuracy, precision, recall, and F1-score.

Notice how the numbers change as you modify the data.

Need Extra Challenge?#

Try building a function that reports all the evaluation metrics for any classification model and dataset.

Bonus: Visualize the confusion matrix as a heatmap.

Recap#

Today you learned about accuracy, precision, recall, F1-score, and the confusion matrix. You also practiced with two real-world datasets.

These skills help you know when your model is really working. Keep practicing!

Thank you for learning with us!#

Like, subscribe, and share if you want more beginner-friendly Python and data mining tutorials.

Try out what you learnedsee you in the next video!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.