Lesson 15 · Python For Machine Learning
Building a Heart Disease Classification Model in Python: Step-by-Step Machine Learning Guide
In this beginner lesson, we will explore how to use Python to analyze and predict heart disease risk. You will learn about variables, data types, and how to…
- CoursePython For Machine Learning
- Lesson15 of 16
- Video13 min
- FormatJupyter notebook · 20 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynb
Welcome to Heart Disease Classification in Python!#
In this beginner lesson, we will explore how to use Python to analyze and predict heart disease risk.
You will learn about variables, data types, and how to work with real-world datasets.
By the end, you will see how these basic building blocks combine in a mini-project.
# Let us start by importing Python's built-in warning filter
import warnings
warnings.filterwarnings('ignore')
What is Classification?#
Classification is a way to sort things into categories.
For example, predicting whether a patient has heart disease is a classification problem.
You will see how we use Python to help solve it step-by-step.
# Let us create our first variable
age = 29
# Variables can store different types of values
name = "Sam"
has_heart_disease = False
Lists and Dictionaries: Storing More Information#
Python offers lists (like a shopping list) and dictionaries (like a contact card).
Dictionaries let us match keys to values, just like name tags with information.
# This dictionary stores one patient's data
patient = {
"name": "Sam",
"age": 29,
"has_heart_disease": False
}
# Access information from a dictionary
print("Name:", patient["name"])
print("Heart Disease:", patient["has_heart_disease"])
# Let us handle a missing key carefully
if "cholesterol" in patient:
print(patient["cholesterol"])
else:
print("No cholesterol data")
Working with Real Data#
Next, we will load a real heart disease dataset to practice our new skills.
We will see how data for hundreds of patients is stored and explored.
# Data setup
import pandas as pd
url = "https://raw.githubusercontent.com/Breezercoder/Heart-disease-prediction/main/heart_disease_dataset.csv"
heart_data = pd.read_csv(url)
print("Rows and columns:", heart_data.shape)
heart_data.head()
# What does the target look like?
print(heart_data["target"].value_counts())
# Let us explore some basic statistics
heart_data.describe()
# Split into train and test sets
from sklearn.model_selection import train_test_split
train, test = train_test_split(heart_data, test_size=0.2, random_state=42)
print("Train size:", len(train), "Test size:", len(test))
Building a Simple Rule#
We can use rules to predict heart disease in an easy way before trying machine learning.
Let us see what happens if we guess 'no' for everyone, and compare it to guessing 'yes' for everyone.
# Everyone predicted as healthy (target 0)
test["prediction"] = 0
accuracy = (test["prediction"] == test["target"]).mean()
print("Accuracy if we guess all are healthy:", accuracy)
# Let us try a smarter rule based on a real feature
test["prediction"] = test["age"].apply(lambda x: 1 if x > 50 else 0)
accuracy = (test["prediction"] == test["target"]).mean()
print("Accuracy guessing heart disease if age > 50:", accuracy)
Using Machine Learning#
We can use a machine learning model to find patterns we might miss.
Let us use a decision tree, which splits data by asking simple yes-or-no questions.
# Decision tree classifier example
from sklearn.tree import DecisionTreeClassifier
features = [col for col in heart_data.columns if col != "target"]
clf = DecisionTreeClassifier(random_state=42)
clf.fit(train[features], train["target"])
test["prediction"] = clf.predict(test[features])
accuracy = (test["prediction"] == test["target"]).mean()
print("Decision tree accuracy:", accuracy)
# Try out your own patient!
user_age = int(input("Enter your age: "))
user_sex = int(input("Enter 1 for male, 0 for female: "))
user_cp = int(input("Chest pain type (0-3): "))
user_trestbps = int(input("Resting blood pressure: "))
user_chol = int(input("Serum cholesterol: "))
user_fbs = int(input("Fasting blood sugar > 120mg/dl? (1 = yes, 0 = no): "))
user_restecg = int(input("Rest ECG result (0-2): "))
user_thalach = int(input("Max heart rate achieved: "))
user_exang = int(input("Exercise induced angina (1 = yes, 0 = no): "))
user_oldpeak = float(input("ST depression: "))
user_slope = int(input("Slope of ST segment (0-2): "))
user_ca = int(input("Number of vessels colored (0-3): "))
user_thal = int(input("Thalassemia (1 = normal, 2 = fixed defect, 3 = reversible defect): "))
user_row = [[user_age, user_sex, user_cp, user_trestbps, user_chol, user_fbs, user_restecg, user_thalach, user_exang, user_oldpeak, user_slope, user_ca, user_thal]]
user_pred = clf.predict(user_row)[0]
if user_pred == 1:
print("Prediction: Risk of Heart Disease")
else:
print("Prediction: No Heart Disease Detected")
# Best practices: check your work
missing = heart_data.isnull().sum()
print("Missing values in each column:\n", missing)
# Common mistakes: wrong data shapes
try:
clf.predict([[29]])
except Exception as e:
print("Error message:", e)
Extra Tips#
- Always start small and test each part.
- Print often to check your progress.
- Try swapping methods and comparing results.
# Challenge: Change the tree depth and see the result!
new_clf = DecisionTreeClassifier(max_depth=3, random_state=42)
new_clf.fit(train[features], train["target"])
score = new_clf.score(test[features], test["target"])
print("Tree with max_depth=3 has accuracy:", score)
# Challenge: Predict a patient using only age and cholesterol
short_features = ["age", "chol"]
clf_small = DecisionTreeClassifier(random_state=42)
clf_small.fit(train[short_features], train["target"])
small_pred = clf_small.predict(test[short_features])
print("Accuracy with just age and cholesterol:", (small_pred == test["target"]).mean())
Recap: What Have We Learned?#
You can:
- Make, read, and update Python variables
- Work with data dictionaries and lists
- Explore real datasets using pandas
- Test prediction rules and a decision tree model
- Try your own patient example
Every project starts with simple steps. Keep going!
Next Steps & Thank You#
Try rewriting parts of this notebook with your own data.
Practice with different datasets to get comfortable.
If you learned something today, like and subscribe on YouTube for more fun Python lessons!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



