Mathew K Analytics

Lesson 29 · Data Science Projects

Predicting Heart Disease with Logistic Regression: A Step-by-Step Python Guide

This beginner friendly lesson explores how to build a simple machine learning model to predict heart disease. You will learn step by step data loading,…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Heart Disease Prediction Using Logistic Regression#

  • This beginner friendly lesson explores how to build a simple machine learning model to predict heart disease.

  • You will learn step by step data loading, preprocessing, visualization, model training, and evaluation.

  • We use real world health data for an inspiring hands on project.

  • Predicting heart disease risk can help save lives by identifying high risk patients early.

  • By the end, you will understand the basics of classification and how logistic regression works.

  • Subscribe to the channel for more helpful data science tutorials!

# Always suppress warnings for a cleaner output.
import warnings; warnings.filterwarnings("ignore")
import numpy as np
np.random.seed(42)

Data setup#

  • We use a real heart disease dataset with various medical features.
  • Each row is a patient record. There is a diagnosis column with 1 for disease and 0 for no disease.
  • Exploring health data can teach us a lot about risk factors.
  • Let us preview the data to understand what we are working with.
# Data setup: Load Heart Disease Dataset
import kagglehub, os, pandas as pd
path = kagglehub.dataset_download('johnsmith88/heart-disease-dataset')
files = os.listdir(path)
df = pd.read_csv(os.path.join(path, files[0]))
print(df.shape)
print(df.head(3))
(1025, 14)
   age  sex  cp  trestbps  chol  fbs  restecg  thalach  exang  oldpeak  slope  \
0   52    1   0       125   212    0        1      168      0      1.0      2   
1   53    1   0       140   203    1        0      155      1      3.1      0   
2   70    1   0       145   174    0        1      125      1      2.6      0   

   ca  thal  target  
0   2     3       0  
1   0     3       0  
2   0     3       0  
# Check for missing values
df.isnull().sum()
age         0
sex         0
cp          0
trestbps    0
chol        0
fbs         0
restecg     0
thalach     0
exang       0
oldpeak     0
slope       0
ca          0
thal        0
target      0
dtype: int64
# Basic data info
df.info()
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 1025 entries, 0 to 1024
Data columns (total 14 columns):
 #   Column    Non-Null Count  Dtype  
---  ------    --------------  -----  
 0   age       1025 non-null   int64  
 1   sex       1025 non-null   int64  
 2   cp        1025 non-null   int64  
 3   trestbps  1025 non-null   int64  
 4   chol      1025 non-null   int64  
 5   fbs       1025 non-null   int64  
 6   restecg   1025 non-null   int64  
 7   thalach   1025 non-null   int64  
 8   exang     1025 non-null   int64  
 9   oldpeak   1025 non-null   float64
 10  slope     1025 non-null   int64  
 11  ca        1025 non-null   int64  
 12  thal      1025 non-null   int64  
 13  target    1025 non-null   int64  
dtypes: float64(1), int64(13)
memory usage: 112.2 KB
# View summary statistics
df.describe()
age sex cp trestbps chol fbs restecg thalach exang oldpeak slope ca thal target
count 1025.000000 1025.000000 1025.000000 1025.000000 1025.00000 1025.000000 1025.000000 1025.000000 1025.000000 1025.000000 1025.000000 1025.000000 1025.000000 1025.000000
mean 54.434146 0.695610 0.942439 131.611707 246.00000 0.149268 0.529756 149.114146 0.336585 1.071512 1.385366 0.754146 2.323902 0.513171
std 9.072290 0.460373 1.029641 17.516718 51.59251 0.356527 0.527878 23.005724 0.472772 1.175053 0.617755 1.030798 0.620660 0.500070
min 29.000000 0.000000 0.000000 94.000000 126.00000 0.000000 0.000000 71.000000 0.000000 0.000000 0.000000 0.000000 0.000000 0.000000
25% 48.000000 0.000000 0.000000 120.000000 211.00000 0.000000 0.000000 132.000000 0.000000 0.000000 1.000000 0.000000 2.000000 0.000000
50% 56.000000 1.000000 1.000000 130.000000 240.00000 0.000000 1.000000 152.000000 0.000000 0.800000 1.000000 0.000000 2.000000 1.000000
75% 61.000000 1.000000 2.000000 140.000000 275.00000 0.000000 1.000000 166.000000 1.000000 1.800000 2.000000 1.000000 3.000000 1.000000
max 77.000000 1.000000 3.000000 200.000000 564.00000 1.000000 2.000000 202.000000 1.000000 6.200000 2.000000 4.000000 3.000000 1.000000

Exploratory Data Analysis#

  • Good data analysis means plotting and looking for patterns.
  • Visualizing the diagnosis target is a first step. Are there more healthy or unhealthy patients?
  • Next, we can plot some features to see how they relate to disease presence.
# Plot target variable counts
import matplotlib.pyplot as plt
import seaborn as sns
sns.countplot(data=df, x='target')
plt.title('Count of patients with and without heart disease')
plt.xlabel('Diagnosis (1 = Disease, 0 = No Disease)')
plt.ylabel('Number of patients')
plt.show()
No description has been provided for this image
# Visualize a key numeric feature
sns.histplot(data=df, x='age', hue='target', kde=True, bins=20)
plt.title('Age Distribution by Diagnosis')
plt.xlabel('Age')
plt.ylabel('Number of patients')
plt.legend(['No Disease', 'Disease'], title='Diagnosis')
plt.show()
No description has been provided for this image
# Look for correlation between features
corr = df.corr()
plt.figure(figsize=(10, 8))
sns.heatmap(corr, annot=True, cmap='coolwarm', fmt='.2f')
plt.title('Feature Correlations')
plt.show()
No description has been provided for this image

Data Preparation for Modeling#

  • Before training, we need to split our data into input features and target labels.
  • We also divide the dataset so our model can learn from a part, and then be tested on unseen data.
  • For fair evaluation, let us use a train test split and random_state for repeatable results.
# Split into features and target, and then train test sets
from sklearn.model_selection import train_test_split
X = df.drop('target', axis=1)
y = df['target']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)
print('Train samples:', X_train.shape[0])
print('Test samples:', X_test.shape[0])
Train samples: 717
Test samples: 308
# Scale data to standard range
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)
X_test_scaled = scaler.transform(X_test)

Train Logistic Regression Model#

  • Logistic regression is a simple but powerful model for binary classification.
  • It predicts the odds of a certain outcome, here disease or no disease.
  • We first train or fit the model to the training data.
# Initialize and train classifier
from sklearn.linear_model import LogisticRegression
model = LogisticRegression(random_state=42)
model.fit(X_train_scaled, y_train)
LogisticRegression(random_state=42)
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
# Predict on test data
y_pred = model.predict(X_test_scaled)
# Measure model accuracy
from sklearn.metrics import accuracy_score
acc = accuracy_score(y_test, y_pred)
print(f"Test accuracy: {acc:.2f}")
Test accuracy: 0.81
# Show confusion matrix
from sklearn.metrics import confusion_matrix, ConfusionMatrixDisplay
cm = confusion_matrix(y_test, y_pred)
disp = ConfusionMatrixDisplay(confusion_matrix=cm, display_labels=['No Disease', 'Disease'])
disp.plot(cmap='Blues')
plt.title('Confusion Matrix')
plt.show()
No description has been provided for this image
# Print classification report
from sklearn.metrics import classification_report
print(classification_report(y_test, y_pred, target_names=['No Disease', 'Disease']))
              precision    recall  f1-score   support

  No Disease       0.86      0.75      0.80       159
     Disease       0.76      0.87      0.81       149

    accuracy                           0.81       308
   macro avg       0.81      0.81      0.80       308
weighted avg       0.81      0.81      0.80       308

Try a Prediction Yourself#

  • You can enter patient features to estimate disease risk using this model.
  • Let us use a test sample to see what prediction is made.
# Enter new patient data for prediction
sample_idx = int(input("Enter a test sample index from 0 to {}: ".format(X_test.shape[0] - 1)))
sample_features = X_test.iloc[sample_idx]
sample_scaled = scaler.transform([sample_features])[0]
pred = model.predict([sample_scaled])[0]
print("Prediction for this patient:", "Disease" if pred == 1 else "No Disease")
print("Actual diagnosis:", "Disease" if y_test.iloc[sample_idx] == 1 else "No Disease")
Prediction for this patient: Disease
Actual diagnosis: Disease

What Did We Learn?#

  • You imported health data, explored important features, and practiced visualization.

  • You prepared the data so the model could learn patterns safely and accurately.

  • You trained a logistic regression classifier to predict if a patient has heart disease.

  • The models results can be measured by accuracy, confusion matrix, and precision recall.

  • You finally tried interactive prediction and model checking with new examples.

  • This workflow is common for any binary classification in healthcare or other fields.

  • If you want more lessons, please like and subscribe for future tutorials!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.