Lesson 14 · Python Fundamentals
Understanding Accuracy, Precision, Recall, F1 Score, and ROC Curve for Model Evaluation in Python
In this lesson, we will learn the basics of measuring how well a machine learning model works. We will use simple, real-world data and show you how to…
- CoursePython Fundamentals
- Lesson14 of 22
- Video15 min
- FormatJupyter notebook · 15 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynb
Python Evaluation Metrics: Accuracy, Precision, Recall, F1, ROC#
In this lesson, we will learn the basics of measuring how well a machine learning model works. We will use simple, real-world data and show you how to calculate the most important metrics: Accuracy, Precision, Recall, F1 Score, and ROC curves.
By the end, you will know how to check your models and why these metrics matter.
# Setup - Filter warnings for a clean experience
import warnings
warnings.filterwarnings('ignore') # Hide warnings so output stays simple
What does model evaluation mean?#
When we make a machine learning model, we do not just want to make predictions. We want to check how good those predictions are.
Evaluation metrics are numbers that help us measure how well a model is making predictions.
Let us see how to do this in Python.
# Data setup: Loading Titanic dataset
import pandas as pd
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
df = pd.read_csv(url)
print('Shape:', df.shape)
df.head()
# Check the survival counts (our target variable)
df['Survived'].value_counts()
# Prepare data: Choose simple features and split into train/test sets
from sklearn.model_selection import train_test_split
X = df[['Pclass', 'Age', 'SibSp', 'Fare']].copy()
X['Age'].fillna(df['Age'].median(), inplace=True)
y = df['Survived']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)
print(X_train.shape, X_test.shape)
# Build a simple logistic regression model
from sklearn.linear_model import LogisticRegression
model = LogisticRegression(max_iter=200)
model.fit(X_train, y_train)
# Predict survival on the test set
y_pred = model.predict(X_test)
print('First 10 predictions:', y_pred[:10])
Understanding accuracy#
- Accuracy is simply the percent of guesses the model got right.
High accuracy is good, but sometimes it is not the whole story.
# Calculate accuracy score
from sklearn.metrics import accuracy_score
acc = accuracy_score(y_test, y_pred)
print('Accuracy:', acc)
What about precision?#
- Precision measures how many positive predictions were actually correct.
It tells us: When the model says 'yes', how often is it right?
# Calculate precision score
from sklearn.metrics import precision_score
prec = precision_score(y_test, y_pred)
print('Precision:', prec)
What does recall mean?#
- Recall shows what fraction of true positives were found.
It answers: Out of all the real 'yes' cases, how many did our model catch?
# Calculate recall score
from sklearn.metrics import recall_score
rec = recall_score(y_test, y_pred)
print('Recall:', rec)
F1 Score: The balance between precision and recall#
- F1 Score is a single number that balances precision and recall.
Useful when you care about both, not just one.
# Calculate F1 score
from sklearn.metrics import f1_score
f1 = f1_score(y_test, y_pred)
print('F1 Score:', f1)
Bonus: See all metrics at once with a classification report#
Scikit-learn can show you a report with accuracy, precision, recall, and F1 score in one table.
Very helpful for quick model checks.
# Print the full classification report
from sklearn.metrics import classification_report
report = classification_report(y_test, y_pred, target_names=['Did Not Survive','Survived'])
print(report)
ROC Curve: Measuring separability#
- ROC stands for Receiver Operating Characteristic.
- It shows how well your model separates the classes, at all possible thresholds.
A perfect model makes a big elbow at the top left of the ROC plot.
# Draw the ROC Curve
from sklearn.metrics import roc_curve, auc
import matplotlib.pyplot as plt
y_prob = model.predict_proba(X_test)[:, 1]
fpr, tpr, thresholds = roc_curve(y_test, y_prob)
roc_auc = auc(fpr, tpr)
plt.figure()
plt.plot(fpr, tpr, color='blue', lw=2, label='ROC curve (area = %0.2f)' % roc_auc)
plt.plot([0, 1], [0, 1], color='gray', lw=1, linestyle='--')
plt.xlabel('False Positive Rate')
plt.ylabel('True Positive Rate')
plt.title('Receiver Operating Characteristic')
plt.legend(loc='lower right')
plt.show()
# Find the best threshold for classifying as survived
import numpy as np
for thresh in np.arange(0, 1.01, 0.1):
preds = (y_prob >= thresh).astype(int)
print('Threshold:', round(thresh,2), 'Precision:', round(precision_score(y_test, preds),2), 'Recall:', round(recall_score(y_test, preds),2))
Real-world tips for metric choice#
- Accuracy can be misleading if there are a lot more of one label than the other.
- Precision is important if false alarms are bad.
- Recall matters if missing real cases is a big problem.
- F1 Score is great when you care about both.
- ROC curve is good for checking different cutoff points.
Think about your problem before picking a metric!
# Challenge: Try changing one model input
X_train2 = X_train.copy()
X_train2['Fare'] = X_train2['Fare'] * 2
model2 = LogisticRegression(max_iter=200)
model2.fit(X_train2, y_train)
y_pred2 = model2.predict(X_test)
print('Accuracy with fare doubled:', accuracy_score(y_test, y_pred2))
Recap: What did you learn?#
You learned how to:
- Set up data and make a simple model
- Measure performance with accuracy, precision, recall, F1, and ROC curve
- Understand which metric to use and why
- Try experiments to see how changes affect metrics
You are now ready to use these metrics in real projects!
Thanks for learning with us!#
If this helped, please subscribe to our channel for more free Python and data videos.
Leave your questions or ideas in the comments below.
Happy coding!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



