Lesson 79 · Data Science Projects
Applying Binary Classification for Accurate Diabetes Prediction in Healthcare
Welcome to Data Mining with a real healthcare dataset! In this session, we will use the Pima Indians Diabetes dataset to predict diabetes using binary…
- CourseData Science Projects
- Lesson79 of 33
- Video17 min
- FormatJupyter notebook · 15 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbWeek 6: Healthcare Data Mining Diabetes Prediction with Python#
Welcome to Data Mining with a real healthcare dataset!
In this session, we will use the Pima Indians Diabetes dataset to predict diabetes using binary classification.
You will practice loading data, exploring and cleaning it, building models, and checking their accuracy.
Data mining helps discover useful patterns from health records. This guides decisions about care, risk, and prevention.
Let us dive in and start working with data!
# Suppress all warnings to keep our workspace clean
import warnings; warnings.filterwarnings('ignore')
import numpy as np
np.random.seed(42)
# Data setup (Pima Indians Diabetes Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/jbrownlee/Datasets/master/pima-indians-diabetes.data.csv'
cols = ['Pregnancies','Glucose','BloodPressure','SkinThickness','Insulin','BMI','DiabetesPedigree','Age','Outcome']
df = pd.read_csv(url, names=cols)
print(df.shape)
print(df.head(3))
What does the dataset look like?#
Each row describes one patient. The last column, Outcome, is 1 if the patient has diabetes, 0 if not.
Let us see how much data we have, and how to find useful insights!
# How many people have diabetes in our dataset?
print(df['Outcome'].value_counts())
# Check for missing values (sometimes zeros are used instead of NaN)
print((df==0).sum())
# Count zero values in Glucose, BloodPressure, SkinThickness, Insulin, BMI
zero_cols = ['Glucose','BloodPressure','SkinThickness','Insulin','BMI']
for col in zero_cols:
n_zeros = (df[col]==0).sum()
print(f"{col}: {n_zeros} zeros")
Data Cleaning: Fixing Zeros#
For our analysis, we will replace zeros in Glucose, BloodPressure, SkinThickness, Insulin, and BMI with the column median.
# Replace unrealistic zeros with the median in selected columns
for col in zero_cols:
median = df[df[col]!=0][col].median()
df.loc[df[col]==0, col] = median
print(df[zero_cols].head(3))
Quick Exploration: Statistics and Correlations#
Basic statistics and correlation checks help us spot important patterns.
For example, does higher glucose mean higher diabetes risk?
# Summary statistics for all columns
print(df.describe())
# Correlation matrix shows relationships between variables
print(df.corr()['Outcome'].sort_values(ascending=False))
# Let us plot histograms for features using pandas
import matplotlib.pyplot as plt
df.hist(bins=20, figsize=(12,8))
plt.tight_layout()
plt.show()
# Create X (inputs) and y (target)
X = df.drop('Outcome', axis=1)
y = df['Outcome']
# Split into training and test data
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
print(X_train.shape, X_test.shape)
Model Time: Logistic Regression#
Logistic regression helps predict binary outcomes in this case, diabetes Yes/No.
# Fit a logistic regression model
from sklearn.linear_model import LogisticRegression
clf = LogisticRegression(max_iter=200)
clf.fit(X_train, y_train)
# Predict on test data and check accuracy
y_pred = clf.predict(X_test)
from sklearn.metrics import accuracy_score
acc = accuracy_score(y_test, y_pred)
print(f'Accuracy: {acc:.2f}')
# Show confusion matrix for more detail
from sklearn.metrics import confusion_matrix
cm = confusion_matrix(y_test, y_pred)
print(cm)
# Try predicting diabetes for a single new input
import numpy as np
new_patient = np.array([[2, 120, 68, 25, 100, 28.5, 0.33, 35]])
pred = clf.predict(new_patient)
print('Will this patient likely develop diabetes?', 'Yes' if pred[0]==1 else 'No')
Congratulations!#
You have cleaned, explored, and built a predictive diabetes model in Python.
Try changing parameters or models to see what improves results.
Data mining can help spot health risks, guide care, and save lives.
Ready for more? Try another health dataset or add more features!
Keep practicing and share your progress by commenting below!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



