Mathew K Analytics

Lesson 14 · Python Fundamentals

Understanding Accuracy, Precision, Recall, F1 Score, and ROC Curve for Model Evaluation in Python

In this lesson, we will learn the basics of measuring how well a machine learning model works. We will use simple, real-world data and show you how to…

⬇ Download notebookOpen in Colab ↗

📓 Full notebook

Download .ipynb
 

Python Evaluation Metrics: Accuracy, Precision, Recall, F1, ROC#

In this lesson, we will learn the basics of measuring how well a machine learning model works. We will use simple, real-world data and show you how to calculate the most important metrics: Accuracy, Precision, Recall, F1 Score, and ROC curves.

By the end, you will know how to check your models and why these metrics matter.

# Setup - Filter warnings for a clean experience
import warnings
warnings.filterwarnings('ignore')  # Hide warnings so output stays simple

What does model evaluation mean?#

When we make a machine learning model, we do not just want to make predictions. We want to check how good those predictions are.

Evaluation metrics are numbers that help us measure how well a model is making predictions.

Let us see how to do this in Python.

# Data setup: Loading Titanic dataset
import pandas as pd
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
df = pd.read_csv(url)
print('Shape:', df.shape)
df.head()
Shape: (891, 12)
PassengerId Survived Pclass Name Sex Age SibSp Parch Ticket Fare Cabin Embarked
0 1 0 3 Braund, Mr. Owen Harris male 22.0 1 0 A/5 21171 7.2500 NaN S
1 2 1 1 Cumings, Mrs. John Bradley (Florence Briggs Th... female 38.0 1 0 PC 17599 71.2833 C85 C
2 3 1 3 Heikkinen, Miss. Laina female 26.0 0 0 STON/O2. 3101282 7.9250 NaN S
3 4 1 1 Futrelle, Mrs. Jacques Heath (Lily May Peel) female 35.0 1 0 113803 53.1000 C123 S
4 5 0 3 Allen, Mr. William Henry male 35.0 0 0 373450 8.0500 NaN S
# Check the survival counts (our target variable)
df['Survived'].value_counts()
Survived
0    549
1    342
Name: count, dtype: int64
# Prepare data: Choose simple features and split into train/test sets
from sklearn.model_selection import train_test_split
X = df[['Pclass', 'Age', 'SibSp', 'Fare']].copy()
X['Age'].fillna(df['Age'].median(), inplace=True)
y = df['Survived']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)
print(X_train.shape, X_test.shape)
(623, 4) (268, 4)
# Build a simple logistic regression model
from sklearn.linear_model import LogisticRegression
model = LogisticRegression(max_iter=200)
model.fit(X_train, y_train)
LogisticRegression(max_iter=200)
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
# Predict survival on the test set
y_pred = model.predict(X_test)
print('First 10 predictions:', y_pred[:10])
First 10 predictions: [0 0 0 1 0 1 0 0 0 1]

Understanding accuracy#

  • Accuracy is simply the percent of guesses the model got right.

High accuracy is good, but sometimes it is not the whole story.

# Calculate accuracy score
from sklearn.metrics import accuracy_score
acc = accuracy_score(y_test, y_pred)
print('Accuracy:', acc)
Accuracy: 0.7201492537313433

What about precision?#

  • Precision measures how many positive predictions were actually correct.

It tells us: When the model says 'yes', how often is it right?

# Calculate precision score
from sklearn.metrics import precision_score
prec = precision_score(y_test, y_pred)
print('Precision:', prec)
Precision: 0.78125

What does recall mean?#

  • Recall shows what fraction of true positives were found.

It answers: Out of all the real 'yes' cases, how many did our model catch?

# Calculate recall score
from sklearn.metrics import recall_score
rec = recall_score(y_test, y_pred)
print('Recall:', rec)
Recall: 0.45045045045045046

F1 Score: The balance between precision and recall#

  • F1 Score is a single number that balances precision and recall.

Useful when you care about both, not just one.

# Calculate F1 score
from sklearn.metrics import f1_score
f1 = f1_score(y_test, y_pred)
print('F1 Score:', f1)
F1 Score: 0.5714285714285714

Bonus: See all metrics at once with a classification report#

Scikit-learn can show you a report with accuracy, precision, recall, and F1 score in one table.

Very helpful for quick model checks.

# Print the full classification report
from sklearn.metrics import classification_report
report = classification_report(y_test, y_pred, target_names=['Did Not Survive','Survived'])
print(report)
                 precision    recall  f1-score   support

Did Not Survive       0.70      0.91      0.79       157
       Survived       0.78      0.45      0.57       111

       accuracy                           0.72       268
      macro avg       0.74      0.68      0.68       268
   weighted avg       0.73      0.72      0.70       268

ROC Curve: Measuring separability#

  • ROC stands for Receiver Operating Characteristic.
  • It shows how well your model separates the classes, at all possible thresholds.

A perfect model makes a big elbow at the top left of the ROC plot.

# Draw the ROC Curve
from sklearn.metrics import roc_curve, auc
import matplotlib.pyplot as plt
y_prob = model.predict_proba(X_test)[:, 1]
fpr, tpr, thresholds = roc_curve(y_test, y_prob)
roc_auc = auc(fpr, tpr)
plt.figure()
plt.plot(fpr, tpr, color='blue', lw=2, label='ROC curve (area = %0.2f)' % roc_auc)
plt.plot([0, 1], [0, 1], color='gray', lw=1, linestyle='--')
plt.xlabel('False Positive Rate')
plt.ylabel('True Positive Rate')
plt.title('Receiver Operating Characteristic')
plt.legend(loc='lower right')
plt.show()
No description has been provided for this image
# Find the best threshold for classifying as survived
import numpy as np
for thresh in np.arange(0, 1.01, 0.1):
    preds = (y_prob >= thresh).astype(int)
    print('Threshold:', round(thresh,2), 'Precision:', round(precision_score(y_test, preds),2), 'Recall:', round(recall_score(y_test, preds),2))
    
Threshold: 0.0 Precision: 0.41 Recall: 1.0
Threshold: 0.1 Precision: 0.42 Recall: 0.99
Threshold: 0.2 Precision: 0.44 Recall: 0.97
Threshold: 0.3 Precision: 0.58 Recall: 0.78
Threshold: 0.4 Precision: 0.7 Recall: 0.62
Threshold: 0.5 Precision: 0.78 Recall: 0.45
Threshold: 0.6 Precision: 0.85 Recall: 0.35
Threshold: 0.7 Precision: 0.82 Recall: 0.16
Threshold: 0.8 Precision: 0.67 Recall: 0.05
Threshold: 0.9 Precision: 0.0 Recall: 0.0
Threshold: 1.0 Precision: 0.0 Recall: 0.0

Real-world tips for metric choice#

  • Accuracy can be misleading if there are a lot more of one label than the other.
  • Precision is important if false alarms are bad.
  • Recall matters if missing real cases is a big problem.
  • F1 Score is great when you care about both.
  • ROC curve is good for checking different cutoff points.

Think about your problem before picking a metric!

# Challenge: Try changing one model input
X_train2 = X_train.copy()
X_train2['Fare'] = X_train2['Fare'] * 2
model2 = LogisticRegression(max_iter=200)
model2.fit(X_train2, y_train)
y_pred2 = model2.predict(X_test)
print('Accuracy with fare doubled:', accuracy_score(y_test, y_pred2))
Accuracy with fare doubled: 0.7238805970149254

Recap: What did you learn?#

You learned how to:

  • Set up data and make a simple model
  • Measure performance with accuracy, precision, recall, F1, and ROC curve
  • Understand which metric to use and why
  • Try experiments to see how changes affect metrics

You are now ready to use these metrics in real projects!

Thanks for learning with us!#

If this helped, please subscribe to our channel for more free Python and data videos.

Leave your questions or ideas in the comments below.

Happy coding!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.