Lesson 29 · Data Science Projects
Predicting Heart Disease with Logistic Regression: A Step-by-Step Python Guide
This beginner friendly lesson explores how to build a simple machine learning model to predict heart disease. You will learn step by step data loading,…
- CourseData Science Projects
- Lesson29 of 33
- Video23 min
- FormatJupyter notebook · 16 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbHeart Disease Prediction Using Logistic Regression#
This beginner friendly lesson explores how to build a simple machine learning model to predict heart disease.
You will learn step by step data loading, preprocessing, visualization, model training, and evaluation.
We use real world health data for an inspiring hands on project.
Predicting heart disease risk can help save lives by identifying high risk patients early.
By the end, you will understand the basics of classification and how logistic regression works.
Subscribe to the channel for more helpful data science tutorials!
# Always suppress warnings for a cleaner output.
import warnings; warnings.filterwarnings("ignore")
import numpy as np
np.random.seed(42)
Data setup#
- We use a real heart disease dataset with various medical features.
- Each row is a patient record. There is a diagnosis column with 1 for disease and 0 for no disease.
- Exploring health data can teach us a lot about risk factors.
- Let us preview the data to understand what we are working with.
# Data setup: Load Heart Disease Dataset
import kagglehub, os, pandas as pd
path = kagglehub.dataset_download('johnsmith88/heart-disease-dataset')
files = os.listdir(path)
df = pd.read_csv(os.path.join(path, files[0]))
print(df.shape)
print(df.head(3))
# Check for missing values
df.isnull().sum()
# Basic data info
df.info()
# View summary statistics
df.describe()
Exploratory Data Analysis#
- Good data analysis means plotting and looking for patterns.
- Visualizing the diagnosis target is a first step. Are there more healthy or unhealthy patients?
- Next, we can plot some features to see how they relate to disease presence.
# Plot target variable counts
import matplotlib.pyplot as plt
import seaborn as sns
sns.countplot(data=df, x='target')
plt.title('Count of patients with and without heart disease')
plt.xlabel('Diagnosis (1 = Disease, 0 = No Disease)')
plt.ylabel('Number of patients')
plt.show()
# Visualize a key numeric feature
sns.histplot(data=df, x='age', hue='target', kde=True, bins=20)
plt.title('Age Distribution by Diagnosis')
plt.xlabel('Age')
plt.ylabel('Number of patients')
plt.legend(['No Disease', 'Disease'], title='Diagnosis')
plt.show()
# Look for correlation between features
corr = df.corr()
plt.figure(figsize=(10, 8))
sns.heatmap(corr, annot=True, cmap='coolwarm', fmt='.2f')
plt.title('Feature Correlations')
plt.show()
Data Preparation for Modeling#
- Before training, we need to split our data into input features and target labels.
- We also divide the dataset so our model can learn from a part, and then be tested on unseen data.
- For fair evaluation, let us use a train test split and random_state for repeatable results.
# Split into features and target, and then train test sets
from sklearn.model_selection import train_test_split
X = df.drop('target', axis=1)
y = df['target']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)
print('Train samples:', X_train.shape[0])
print('Test samples:', X_test.shape[0])
# Scale data to standard range
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)
X_test_scaled = scaler.transform(X_test)
Train Logistic Regression Model#
- Logistic regression is a simple but powerful model for binary classification.
- It predicts the odds of a certain outcome, here disease or no disease.
- We first train or fit the model to the training data.
# Initialize and train classifier
from sklearn.linear_model import LogisticRegression
model = LogisticRegression(random_state=42)
model.fit(X_train_scaled, y_train)
# Predict on test data
y_pred = model.predict(X_test_scaled)
# Measure model accuracy
from sklearn.metrics import accuracy_score
acc = accuracy_score(y_test, y_pred)
print(f"Test accuracy: {acc:.2f}")
# Show confusion matrix
from sklearn.metrics import confusion_matrix, ConfusionMatrixDisplay
cm = confusion_matrix(y_test, y_pred)
disp = ConfusionMatrixDisplay(confusion_matrix=cm, display_labels=['No Disease', 'Disease'])
disp.plot(cmap='Blues')
plt.title('Confusion Matrix')
plt.show()
# Print classification report
from sklearn.metrics import classification_report
print(classification_report(y_test, y_pred, target_names=['No Disease', 'Disease']))
Try a Prediction Yourself#
- You can enter patient features to estimate disease risk using this model.
- Let us use a test sample to see what prediction is made.
# Enter new patient data for prediction
sample_idx = int(input("Enter a test sample index from 0 to {}: ".format(X_test.shape[0] - 1)))
sample_features = X_test.iloc[sample_idx]
sample_scaled = scaler.transform([sample_features])[0]
pred = model.predict([sample_scaled])[0]
print("Prediction for this patient:", "Disease" if pred == 1 else "No Disease")
print("Actual diagnosis:", "Disease" if y_test.iloc[sample_idx] == 1 else "No Disease")
What Did We Learn?#
You imported health data, explored important features, and practiced visualization.
You prepared the data so the model could learn patterns safely and accurately.
You trained a logistic regression classifier to predict if a patient has heart disease.
The models results can be measured by accuracy, confusion matrix, and precision recall.
You finally tried interactive prediction and model checking with new examples.
This workflow is common for any binary classification in healthcare or other fields.
If you want more lessons, please like and subscribe for future tutorials!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



