Mathew K Analytics

Lesson 79 · Data Science Projects

Applying Binary Classification for Accurate Diabetes Prediction in Healthcare

Welcome to Data Mining with a real healthcare dataset! In this session, we will use the Pima Indians Diabetes dataset to predict diabetes using binary…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Week 6: Healthcare Data Mining Diabetes Prediction with Python#

Welcome to Data Mining with a real healthcare dataset!

In this session, we will use the Pima Indians Diabetes dataset to predict diabetes using binary classification.

You will practice loading data, exploring and cleaning it, building models, and checking their accuracy.

Data mining helps discover useful patterns from health records. This guides decisions about care, risk, and prevention.

Let us dive in and start working with data!

# Suppress all warnings to keep our workspace clean
import warnings; warnings.filterwarnings('ignore')
import numpy as np
np.random.seed(42)
# Data setup (Pima Indians Diabetes Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/jbrownlee/Datasets/master/pima-indians-diabetes.data.csv'
cols = ['Pregnancies','Glucose','BloodPressure','SkinThickness','Insulin','BMI','DiabetesPedigree','Age','Outcome']
df = pd.read_csv(url, names=cols)
print(df.shape)
print(df.head(3))
(768, 9)
   Pregnancies  Glucose  BloodPressure  SkinThickness  Insulin   BMI  \
0            6      148             72             35        0  33.6   
1            1       85             66             29        0  26.6   
2            8      183             64              0        0  23.3   

   DiabetesPedigree  Age  Outcome  
0             0.627   50        1  
1             0.351   31        0  
2             0.672   32        1  

What does the dataset look like?#

Each row describes one patient. The last column, Outcome, is 1 if the patient has diabetes, 0 if not.

Let us see how much data we have, and how to find useful insights!

# How many people have diabetes in our dataset?
print(df['Outcome'].value_counts())
Outcome
0    500
1    268
Name: count, dtype: int64
# Check for missing values (sometimes zeros are used instead of NaN)
print((df==0).sum())
Pregnancies         111
Glucose               5
BloodPressure        35
SkinThickness       227
Insulin             374
BMI                  11
DiabetesPedigree      0
Age                   0
Outcome             500
dtype: int64
# Count zero values in Glucose, BloodPressure, SkinThickness, Insulin, BMI
zero_cols = ['Glucose','BloodPressure','SkinThickness','Insulin','BMI']
for col in zero_cols:
    n_zeros = (df[col]==0).sum()
    print(f"{col}: {n_zeros} zeros")
    
Glucose: 5 zeros
BloodPressure: 35 zeros
SkinThickness: 227 zeros
Insulin: 374 zeros
BMI: 11 zeros

Data Cleaning: Fixing Zeros#

For our analysis, we will replace zeros in Glucose, BloodPressure, SkinThickness, Insulin, and BMI with the column median.

# Replace unrealistic zeros with the median in selected columns
for col in zero_cols:
    median = df[df[col]!=0][col].median()
    df.loc[df[col]==0, col] = median
print(df[zero_cols].head(3))
   Glucose  BloodPressure  SkinThickness  Insulin   BMI
0      148             72             35      125  33.6
1       85             66             29      125  26.6
2      183             64             29      125  23.3

Quick Exploration: Statistics and Correlations#

Basic statistics and correlation checks help us spot important patterns.

For example, does higher glucose mean higher diabetes risk?

# Summary statistics for all columns
print(df.describe())
       Pregnancies     Glucose  BloodPressure  SkinThickness     Insulin  \
count   768.000000  768.000000     768.000000     768.000000  768.000000   
mean      3.845052  121.656250      72.386719      29.108073  140.671875   
std       3.369578   30.438286      12.096642       8.791221   86.383060   
min       0.000000   44.000000      24.000000       7.000000   14.000000   
25%       1.000000   99.750000      64.000000      25.000000  121.500000   
50%       3.000000  117.000000      72.000000      29.000000  125.000000   
75%       6.000000  140.250000      80.000000      32.000000  127.250000   
max      17.000000  199.000000     122.000000      99.000000  846.000000   

              BMI  DiabetesPedigree         Age     Outcome  
count  768.000000        768.000000  768.000000  768.000000  
mean    32.455208          0.471876   33.240885    0.348958  
std      6.875177          0.331329   11.760232    0.476951  
min     18.200000          0.078000   21.000000    0.000000  
25%     27.500000          0.243750   24.000000    0.000000  
50%     32.300000          0.372500   29.000000    0.000000  
75%     36.600000          0.626250   41.000000    1.000000  
max     67.100000          2.420000   81.000000    1.000000  
# Correlation matrix shows relationships between variables
print(df.corr()['Outcome'].sort_values(ascending=False))
Outcome             1.000000
Glucose             0.492782
BMI                 0.312038
Age                 0.238356
Pregnancies         0.221898
SkinThickness       0.214873
Insulin             0.203790
DiabetesPedigree    0.173844
BloodPressure       0.165723
Name: Outcome, dtype: float64
# Let us plot histograms for features using pandas
import matplotlib.pyplot as plt
df.hist(bins=20, figsize=(12,8))
plt.tight_layout()
plt.show()
No description has been provided for this image
# Create X (inputs) and y (target)
X = df.drop('Outcome', axis=1)
y = df['Outcome']
# Split into training and test data
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
print(X_train.shape, X_test.shape)
(614, 8) (154, 8)

Model Time: Logistic Regression#

Logistic regression helps predict binary outcomes in this case, diabetes Yes/No.

# Fit a logistic regression model
from sklearn.linear_model import LogisticRegression
clf = LogisticRegression(max_iter=200)
clf.fit(X_train, y_train)
LogisticRegression(max_iter=200)
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
# Predict on test data and check accuracy
y_pred = clf.predict(X_test)
from sklearn.metrics import accuracy_score
acc = accuracy_score(y_test, y_pred)
print(f'Accuracy: {acc:.2f}')
Accuracy: 0.75
# Show confusion matrix for more detail
from sklearn.metrics import confusion_matrix
cm = confusion_matrix(y_test, y_pred)
print(cm)
[[82 17]
 [21 34]]
# Try predicting diabetes for a single new input
import numpy as np
new_patient = np.array([[2, 120, 68, 25, 100, 28.5, 0.33, 35]])
pred = clf.predict(new_patient)
print('Will this patient likely develop diabetes?', 'Yes' if pred[0]==1 else 'No')
Will this patient likely develop diabetes? No

Congratulations!#

You have cleaned, explored, and built a predictive diabetes model in Python.

Try changing parameters or models to see what improves results.

Data mining can help spot health risks, guide care, and save lives.

Ready for more? Try another health dataset or add more features!

Keep practicing and share your progress by commenting below!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.