Lesson 20 · Data Mining
Understanding Classifier Evaluation: Accuracy, Precision, Recall, and F1-Score Explained
Welcome to this lesson where we explore how to evaluate machine learning classifiers. We will talk about concepts, write code, and do hands-on exercises.…
- CourseData Mining
- Lesson20 of 31
- Video20 min
- FormatJupyter notebook · 18 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbWeek 56: Evaluating ClassifiersAccuracy, Precision, Recall, and F1-Score#
Welcome to this lesson where we explore how to evaluate machine learning classifiers. We will talk about concepts, write code, and do hands-on exercises.
These metrics help us understand how well our models perform in real life.
We will use the Telecom Customer Churn dataset and the Titanic dataset. Let's start with a quick overview of why evaluation matters!
Why do we evaluate classifiers?#
Building a model is only half the job. Evaluation tells us if predictions are meaningful or misleading.
Wrong metrics can hide problems.
Today we will avoid common pitfalls!
# Suppress warnings for a smoother learning experience
import warnings; warnings.filterwarnings('ignore')
import numpy as np
np.random.seed(42)
# Data setup (Telecom Customer Churn Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/IBM/telco-customer-churn-on-icp4d/master/data/Telco-Customer-Churn.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
# Check for missing values
print(df.isnull().sum().sort_values(ascending=False).head(5))
# Basic cleanup: Drop customerID and rows with missing values
df = df.drop(['customerID'], axis=1)
df = df.dropna()
print(df.shape)
# Convert categorical columns to numbers
for c in df.select_dtypes('object').columns:
df[c] = df[c].astype('category').cat.codes
print(df.head(3))
# Split the data into features and labels
X = df.drop('Churn', axis=1)
y = df['Churn']
# Train-test split
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, random_state=42)
print(X_train.shape, X_test.shape)
# Train a simple classifier: Logistic Regression
from sklearn.linear_model import LogisticRegression
model = LogisticRegression(max_iter=500)
model.fit(X_train, y_train)
# Make predictions on the test set
y_pred = model.predict(X_test)
Understanding Accuracy#
Accuracy is the fraction of correct predictions out of all predictions.
Works well when classes are balanced. Can be misleading for imbalanced data!
Let us compute it next.
# Calculate accuracy
from sklearn.metrics import accuracy_score
accuracy = accuracy_score(y_test, y_pred)
print('Accuracy:', accuracy)
Precision and Recall#
Precision answers: Of all predicted positives, how many were truly positive? Recall answers: Of all actual positives, how many did the model find?
Both matter if making mistakes is expensive!
# Calculate precision and recall
from sklearn.metrics import precision_score, recall_score
precision = precision_score(y_test, y_pred)
recall = recall_score(y_test, y_pred)
print('Precision:', precision)
print('Recall:', recall)
What is F1-Score?#
F1-Score combines precision and recall into one metric. It is their harmonic mean.
Best when you want a balance between both.
# Calculate F1-Score
from sklearn.metrics import f1_score
f1 = f1_score(y_test, y_pred)
print('F1-Score:', f1)
# Confusion matrix for deeper insight
from sklearn.metrics import confusion_matrix
cm = confusion_matrix(y_test, y_pred)
print(cm)
# Show classification report
from sklearn.metrics import classification_report
print(classification_report(y_test, y_pred))
Switching datasets: Titanic#
Let us try a totally different dataset to check what happens when class sizes change.
Ready? Here we go!
# Data setup (Titanic Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
df2 = pd.read_csv(url)
print(df2.shape)
print(df2.head(3))
# Prepare basic Titanic features
df2 = df2[['Survived','Pclass','Sex','Age','SibSp','Fare']]
df2 = df2.dropna()
df2['Sex'] = df2['Sex'].astype('category').cat.codes
X2 = df2.drop('Survived', axis=1)
y2 = df2['Survived']
# Train-test split for Titanic
from sklearn.model_selection import train_test_split
X2_train, X2_test, y2_train, y2_test = train_test_split(X2, y2, test_size=0.2, random_state=42)
# Train and evaluate a logistic regression model for Titanic
from sklearn.linear_model import LogisticRegression
model2 = LogisticRegression(max_iter=300)
model2.fit(X2_train, y2_train)
y2_pred = model2.predict(X2_test)
from sklearn.metrics import classification_report
print(classification_report(y2_test, y2_pred))
Practice and Try It Yourself!#
Switch back to your favorite dataset. Compute accuracy, precision, recall, and F1-score.
Notice how the numbers change as you modify the data.
Need Extra Challenge?#
Try building a function that reports all the evaluation metrics for any classification model and dataset.
Bonus: Visualize the confusion matrix as a heatmap.
Recap#
Today you learned about accuracy, precision, recall, F1-score, and the confusion matrix. You also practiced with two real-world datasets.
These skills help you know when your model is really working. Keep practicing!
Thank you for learning with us!#
Like, subscribe, and share if you want more beginner-friendly Python and data mining tutorials.
Try out what you learnedsee you in the next video!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



