Mathew K Analytics

Lesson 27 · Data Science Projects

Building Accurate Disease Prediction Models with Machine Learning: A Step-by-Step Guide

Learn how data mining helps in predicting diseases. Understand why these models are important in healthcare. We use real medical data for hands-on learning.…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Disease Prediction Using Machine Learning#

  • Learn how data mining helps in predicting diseases.
  • Understand why these models are important in healthcare.
  • We use real medical data for hands-on learning.

What will you learn?#

  • How to load a classic medical dataset for diabetes prediction.
  • Understand how to clean and explore the data.
  • Build a basic machine learning model for disease prediction.
  • Learn to evaluate model accuracy and spot important features.
  • Visualize your results so they are easy to understand.
# Always start by ignoring future warnings
import warnings; warnings.filterwarnings("ignore")
import numpy as np
np.random.seed(42)
 
# Data setup
import pandas as pd
url = 'https://raw.githubusercontent.com/jbrownlee/Datasets/master/pima-indians-diabetes.data.csv'
cols = ['Pregnancies','Glucose','BloodPressure','SkinThickness','Insulin','BMI','DiabetesPedigree','Age','Outcome']
df = pd.read_csv(url, names=cols)
print(df.shape)
print(df.head(3))
(768, 9)
   Pregnancies  Glucose  BloodPressure  SkinThickness  Insulin   BMI  \
0            6      148             72             35        0  33.6   
1            1       85             66             29        0  26.6   
2            8      183             64              0        0  23.3   

   DiabetesPedigree  Age  Outcome  
0             0.627   50        1  
1             0.351   31        0  
2             0.672   32        1  
# What do the columns mean?
for col in df.columns:
    print(col)
Pregnancies
Glucose
BloodPressure
SkinThickness
Insulin
BMI
DiabetesPedigree
Age
Outcome
# Check for missing values
print(df.isnull().sum())
Pregnancies         0
Glucose             0
BloodPressure       0
SkinThickness       0
Insulin             0
BMI                 0
DiabetesPedigree    0
Age                 0
Outcome             0
dtype: int64
# Quick data summary
print(df.describe())
       Pregnancies     Glucose  BloodPressure  SkinThickness     Insulin  \
count   768.000000  768.000000     768.000000     768.000000  768.000000   
mean      3.845052  120.894531      69.105469      20.536458   79.799479   
std       3.369578   31.972618      19.355807      15.952218  115.244002   
min       0.000000    0.000000       0.000000       0.000000    0.000000   
25%       1.000000   99.000000      62.000000       0.000000    0.000000   
50%       3.000000  117.000000      72.000000      23.000000   30.500000   
75%       6.000000  140.250000      80.000000      32.000000  127.250000   
max      17.000000  199.000000     122.000000      99.000000  846.000000   

              BMI  DiabetesPedigree         Age     Outcome  
count  768.000000        768.000000  768.000000  768.000000  
mean    31.992578          0.471876   33.240885    0.348958  
std      7.884160          0.331329   11.760232    0.476951  
min      0.000000          0.078000   21.000000    0.000000  
25%     27.300000          0.243750   24.000000    0.000000  
50%     32.000000          0.372500   29.000000    0.000000  
75%     36.600000          0.626250   41.000000    1.000000  
max     67.100000          2.420000   81.000000    1.000000  
# Check for zeros as missing in medical columns
cols_check = ['Glucose', 'BloodPressure', 'SkinThickness', 'Insulin', 'BMI']
for c in cols_check:
    zero_count = (df[c] == 0).sum()
    print(f"{c}: {zero_count} zeros")
Glucose: 5 zeros
BloodPressure: 35 zeros
SkinThickness: 227 zeros
Insulin: 374 zeros
BMI: 11 zeros
# Replace zeros with column median for selected features
for c in cols_check:
    df[c] = df[c].replace(0, df[c].median())
print(df[cols_check].head(3))
   Glucose  BloodPressure  SkinThickness  Insulin   BMI
0      148             72             35     30.5  33.6
1       85             66             29     30.5  26.6
2      183             64             23     30.5  23.3

Explore patterns in your data#

  • Relationships between features can help us understand risk factors.
  • Visual exploration is a great first step before modeling.
# Plot histograms for features
import matplotlib.pyplot as plt
df.hist(bins=15, figsize=(12,7))
plt.tight_layout()
plt.show()
No description has been provided for this image
# Visualize glucose by outcome using a boxplot
import seaborn as sns
sns.boxplot(data=df, x='Outcome', y='Glucose')
plt.title('Glucose by Diabetes Outcome')
plt.show()
No description has been provided for this image
# See how many have diabetes vs not
sns.countplot(x='Outcome', data=df)
plt.title('Class Balance: Diabetes')
plt.show()
No description has been provided for this image
# Set up features and label arrays
X = df.drop('Outcome', axis=1)
y = df['Outcome']
print(X.shape, y.shape)
(768, 8) (768,)
# Split to train and test sets
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
print(X_train.shape, X_test.shape)
(614, 8) (154, 8)

Let us train a machine learning model#

  • We will use logistic regression for this example.
  • It is a common and interpretable method in medicine.
  • This will help us predict the risk of diabetes from patient records.
# Train logistic regression model
from sklearn.linear_model import LogisticRegression
model = LogisticRegression(max_iter=500, solver='lbfgs')
model.fit(X_train, y_train)
print('Training complete.')
Training complete.
# Predict and check accuracy
from sklearn.metrics import accuracy_score
y_pred = model.predict(X_test)
acc = accuracy_score(y_test, y_pred)
print(f"Test Accuracy: {acc:.2f}")
Test Accuracy: 0.76
# Confusion matrix to see detailed results
from sklearn.metrics import confusion_matrix, ConfusionMatrixDisplay
cm = confusion_matrix(y_test, y_pred)
disp = ConfusionMatrixDisplay(cm, display_labels=['No Diabetes','Diabetes'])
disp.plot()
plt.show()
No description has been provided for this image
# Which features matter most?
import numpy as np
coef = model.coef_[0]
for name, value in zip(X.columns, coef):
    print(f'{name}: {value:.2f}')
Pregnancies: 0.07
Glucose: 0.04
BloodPressure: -0.01
SkinThickness: 0.01
Insulin: -0.00
BMI: 0.11
DiabetesPedigree: 0.59
Age: 0.03
# Try the model on a new patient (input example)
input_vals = []
for c in X.columns:
    val = float(input(f'Enter {c}: '))
    input_vals.append(val)
prediction = model.predict([input_vals])[0]
if prediction == 1:
    print('Prediction: Diabetes present')
else:
    print('Prediction: No diabetes')
Prediction: No diabetes

What did you just build?#

  • You loaded real health data and fixed any data problems.
  • You found patterns using charts and basic statistics.
  • You trained, tested, and checked the results of a machine learning model.
  • You predicted outcomes for example patients.
  • You interpreted which features matter most.
# Stretch: Try a random forest for extra accuracy
from sklearn.ensemble import RandomForestClassifier
rf = RandomForestClassifier(n_estimators=100, random_state=42)
rf.fit(X_train, y_train)
rf_acc = rf.score(X_test, y_test)
print(f"Random Forest Test Accuracy: {rf_acc:.2f}")
Random Forest Test Accuracy: 0.75
# Mini project: Try predicting a new disease outcome!
# 1. Find a different health dataset online.
# 2. Load, clean, and explore it with these same steps.
# 3. Train a classifier, make predictions, and visualize results.
# Share your project with a friend and teach them what you learned.
"If you enjoyed this, subscribe to our channel for more fun projects!"
'If you enjoyed this, subscribe to our channel for more fun projects!'
 

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.