Mathew K Analytics

Lesson 15 · Python For Machine Learning

Building a Heart Disease Classification Model in Python: Step-by-Step Machine Learning Guide

In this beginner lesson, we will explore how to use Python to analyze and predict heart disease risk. You will learn about variables, data types, and how to…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb
 

Welcome to Heart Disease Classification in Python!#

In this beginner lesson, we will explore how to use Python to analyze and predict heart disease risk.

You will learn about variables, data types, and how to work with real-world datasets.

By the end, you will see how these basic building blocks combine in a mini-project.

# Let us start by importing Python's built-in warning filter
import warnings
warnings.filterwarnings('ignore')

What is Classification?#

Classification is a way to sort things into categories.

For example, predicting whether a patient has heart disease is a classification problem.

You will see how we use Python to help solve it step-by-step.

# Let us create our first variable
age = 29
# Variables can store different types of values
name = "Sam"
has_heart_disease = False

Lists and Dictionaries: Storing More Information#

Python offers lists (like a shopping list) and dictionaries (like a contact card).

Dictionaries let us match keys to values, just like name tags with information.

# This dictionary stores one patient's data
patient = {
    "name": "Sam",
    "age": 29,
    "has_heart_disease": False
}
# Access information from a dictionary
print("Name:", patient["name"])
print("Heart Disease:", patient["has_heart_disease"])
Name: Sam
Heart Disease: False
# Let us handle a missing key carefully
if "cholesterol" in patient:
    print(patient["cholesterol"])
else:
    print("No cholesterol data")
    
No cholesterol data

Working with Real Data#

Next, we will load a real heart disease dataset to practice our new skills.

We will see how data for hundreds of patients is stored and explored.

# Data setup
import pandas as pd
url = "https://raw.githubusercontent.com/Breezercoder/Heart-disease-prediction/main/heart_disease_dataset.csv"
heart_data = pd.read_csv(url)
print("Rows and columns:", heart_data.shape)
heart_data.head()
Rows and columns: (303, 14)
age sex cp trestbps chol fbs restecg thalach exang oldpeak slope ca thal target
0 67 1 2 126 458 1 0 144 0 2.3 2 3 1 1
1 57 0 0 158 384 0 1 133 0 6.2 0 2 1 0
2 43 0 3 111 286 0 0 130 0 2.8 1 0 2 0
3 71 1 2 189 515 1 1 149 0 2.1 1 0 2 1
4 36 0 0 142 303 0 0 107 1 3.6 1 2 1 1
# What does the target look like?
print(heart_data["target"].value_counts())
target
0    153
1    150
Name: count, dtype: int64
# Let us explore some basic statistics
heart_data.describe()
age sex cp trestbps chol fbs restecg thalach exang oldpeak slope ca thal target
count 303.000000 303.000000 303.00000 303.000000 303.000000 303.000000 303.000000 303.000000 303.000000 303.000000 303.000000 303.000000 303.000000 303.000000
mean 52.267327 0.544554 1.40264 146.511551 352.894389 0.471947 0.465347 134.610561 0.488449 3.030693 1.000000 1.531353 0.957096 0.495050
std 13.896179 0.498835 1.16925 31.336124 127.705381 0.500038 0.499623 39.731468 0.500693 1.786747 0.805609 1.138507 0.846579 0.500803
min 29.000000 0.000000 0.00000 94.000000 126.000000 0.000000 0.000000 71.000000 0.000000 0.000000 0.000000 0.000000 0.000000 0.000000
25% 40.000000 0.000000 0.00000 119.000000 245.500000 0.000000 0.000000 99.000000 0.000000 1.400000 0.000000 0.000000 0.000000 0.000000
50% 53.000000 1.000000 1.00000 148.000000 354.000000 0.000000 0.000000 132.000000 0.000000 3.000000 1.000000 2.000000 1.000000 0.000000
75% 64.000000 1.000000 2.00000 174.000000 456.000000 1.000000 1.000000 169.000000 1.000000 4.600000 2.000000 3.000000 2.000000 1.000000
max 76.000000 1.000000 3.00000 199.000000 563.000000 1.000000 1.000000 201.000000 1.000000 6.200000 2.000000 3.000000 2.000000 1.000000
# Split into train and test sets
from sklearn.model_selection import train_test_split
train, test = train_test_split(heart_data, test_size=0.2, random_state=42)
print("Train size:", len(train), "Test size:", len(test))
Train size: 242 Test size: 61

Building a Simple Rule#

We can use rules to predict heart disease in an easy way before trying machine learning.

Let us see what happens if we guess 'no' for everyone, and compare it to guessing 'yes' for everyone.

# Everyone predicted as healthy (target 0)
test["prediction"] = 0
accuracy = (test["prediction"] == test["target"]).mean()
print("Accuracy if we guess all are healthy:", accuracy)
Accuracy if we guess all are healthy: 0.45901639344262296
# Let us try a smarter rule based on a real feature
test["prediction"] = test["age"].apply(lambda x: 1 if x > 50 else 0)
accuracy = (test["prediction"] == test["target"]).mean()
print("Accuracy guessing heart disease if age > 50:", accuracy)
Accuracy guessing heart disease if age > 50: 0.4098360655737705

Using Machine Learning#

We can use a machine learning model to find patterns we might miss.

Let us use a decision tree, which splits data by asking simple yes-or-no questions.

# Decision tree classifier example
from sklearn.tree import DecisionTreeClassifier
features = [col for col in heart_data.columns if col != "target"]
clf = DecisionTreeClassifier(random_state=42)
clf.fit(train[features], train["target"])
test["prediction"] = clf.predict(test[features])
accuracy = (test["prediction"] == test["target"]).mean()
print("Decision tree accuracy:", accuracy)
Decision tree accuracy: 0.6065573770491803
# Try out your own patient!
user_age = int(input("Enter your age: "))
user_sex = int(input("Enter 1 for male, 0 for female: "))
user_cp = int(input("Chest pain type (0-3): "))
user_trestbps = int(input("Resting blood pressure: "))
user_chol = int(input("Serum cholesterol: "))
user_fbs = int(input("Fasting blood sugar > 120mg/dl? (1 = yes, 0 = no): "))
user_restecg = int(input("Rest ECG result (0-2): "))
user_thalach = int(input("Max heart rate achieved: "))
user_exang = int(input("Exercise induced angina (1 = yes, 0 = no): "))
user_oldpeak = float(input("ST depression: "))
user_slope = int(input("Slope of ST segment (0-2): "))
user_ca = int(input("Number of vessels colored (0-3): "))
user_thal = int(input("Thalassemia (1 = normal, 2 = fixed defect, 3 = reversible defect): "))
user_row = [[user_age, user_sex, user_cp, user_trestbps, user_chol, user_fbs, user_restecg, user_thalach, user_exang, user_oldpeak, user_slope, user_ca, user_thal]]
user_pred = clf.predict(user_row)[0]
if user_pred == 1:
    print("Prediction: Risk of Heart Disease")
else:
    print("Prediction: No Heart Disease Detected")
    
Prediction: Risk of Heart Disease
 
# Best practices: check your work
missing = heart_data.isnull().sum()
print("Missing values in each column:\n", missing)
Missing values in each column:
 age         0
sex         0
cp          0
trestbps    0
chol        0
fbs         0
restecg     0
thalach     0
exang       0
oldpeak     0
slope       0
ca          0
thal        0
target      0
dtype: int64
# Common mistakes: wrong data shapes
try:
    clf.predict([[29]])
except Exception as e:
    print("Error message:", e)
    
Error message: X has 1 features, but DecisionTreeClassifier is expecting 13 features as input.

Extra Tips#

  • Always start small and test each part.
  • Print often to check your progress.
  • Try swapping methods and comparing results.
# Challenge: Change the tree depth and see the result!
new_clf = DecisionTreeClassifier(max_depth=3, random_state=42)
new_clf.fit(train[features], train["target"])
score = new_clf.score(test[features], test["target"])
print("Tree with max_depth=3 has accuracy:", score)
Tree with max_depth=3 has accuracy: 0.5081967213114754
# Challenge: Predict a patient using only age and cholesterol
short_features = ["age", "chol"]
clf_small = DecisionTreeClassifier(random_state=42)
clf_small.fit(train[short_features], train["target"])
small_pred = clf_small.predict(test[short_features])
print("Accuracy with just age and cholesterol:", (small_pred == test["target"]).mean())
Accuracy with just age and cholesterol: 0.5573770491803278

Recap: What Have We Learned?#

You can:

  • Make, read, and update Python variables
  • Work with data dictionaries and lists
  • Explore real datasets using pandas
  • Test prediction rules and a decision tree model
  • Try your own patient example

Every project starts with simple steps. Keep going!

Next Steps & Thank You#

Try rewriting parts of this notebook with your own data.

Practice with different datasets to get comfortable.

If you learned something today, like and subscribe on YouTube for more fun Python lessons!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.