Lesson 31 · Real-World Data Analytics
Python Data Analytics #31: Real Diabetes Risk Factor Analysis in Python
And the start of a brand new domain: healthcare and public health. A real, genuine patient dataset used to hunt for the real factors most strongly tied to a…
- CourseReal-World Data Analytics
- Lesson31 of 100
- FormatJupyter notebook · 30 code cells
- Data1 dataset
What you'll learn
Datasets used in this lesson
Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.
- diabetes_raw.csv preview22.6 KB
📓 Full notebook
Download .ipynbData Analytics 100, Video 31: Real Diabetes Risk Factor Analysis#
- Video thirty-one of the hundred-video real-world data analytics series, and the start of a brand new domain: healthcare and public health.
- A real, genuine patient dataset used to hunt for the real factors most strongly tied to a real diabetes diagnosis.
- Let's get into it.
This lesson is for learners who know basic pandas and want to apply it to health data for the first time. The question is a practical one: in a real patient dataset, which measurements are most strongly linked to a diabetes diagnosis, and how much does risk change as those measurements rise?
The data is a well-known medical dataset of 768 patients loaded from diabetes_raw.csv. Each row is one patient, with clinical measurements such as Glucose, BloodPressure, SkinThickness, Insulin, BMI, DiabetesPedigreeFunction (a score summarising family history of diabetes) and Age, plus the number of Pregnancies. The target column is Outcome, where 1 means the patient was diagnosed with diabetes and 0 means not.
In this lesson you will:
- spot zeros that are really missing values and convert them safely
- compare patients with and without diabetes using
groupbymeans and correlations - group continuous measurements into clinical bands with
pd.cut - score a simple threshold rule using precision and recall
Prerequisites: you should be comfortable with groupby, boolean filtering and basic matplotlib charts. If not, see the earlier lessons in the series.
Part 1: A Real, Famous Medical Dataset#
Before any analysis you need to know how much data you have. The code imports pandas, numpy and matplotlib, reads diabetes_raw.csv with pd.read_csv, and checks .shape, which returns a (rows, columns) pair. The output is (768, 9): 768 patients and 9 columns. That is small enough to reason about row by row, but it also means a few dozen bad values can shift averages noticeably, which matters later in this lesson.
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
diabetes = pd.read_csv('diabetes_raw.csv')
diabetes.shape
Part 2: Real Data Preview#
A quick preview tells you what each column is called and what it holds before you write any calculations. The code calls diabetes.head(5) and then diabetes.columns.tolist(), but Jupyter only displays the last expression, so the output you see is the list of column names. It shows eight measurements (Pregnancies, Glucose, BloodPressure, SkinThickness, Insulin, BMI, DiabetesPedigreeFunction, Age) and the target Outcome. Run head(5) on its own line if you want to see sample rows too.
diabetes.head(5)
diabetes.columns.tolist()
Part 3: Real Outcome Variable#
Every comparison in this lesson is measured against the overall diagnosis rate, so you need it first. value_counts() counts patients in each Outcome class, but only the last expression is shown. Because Outcome is coded 0 and 1, .mean() gives the share of patients equal to 1, and multiplying by 100 and rounding gives a percentage. The output is 34.9: about a third of patients in this dataset have diabetes. Keep this baseline in mind when you judge whether a group's rate is high or low.
outcome_counts = diabetes['Outcome'].value_counts()
outcome_counts
outcome_rate = round(diabetes['Outcome'].mean() * 100, 1)
outcome_rate
Part 4: A Real Data Quality Trap#
Some clinical measurements cannot physically be zero. A living patient does not have a glucose level, blood pressure or BMI of 0, so a zero here really means the value was not recorded. The code compares five columns to 0 with == 0, which produces True or False for each cell, and .sum() counts the Trues per column. The output shows 5 zeros in Glucose, 35 in BloodPressure, 11 in BMI, 227 in SkinThickness and 374 in Insulin. Left alone, these zeros would drag every average down without any warning.
check_cols = ['Glucose', 'BloodPressure', 'SkinThickness', 'Insulin', 'BMI']
zero_counts = (diabetes[check_cols] == 0).sum()
zero_counts
Part 5: Real Cleaning, Zeros to Missing#
Once you know the zeros are placeholders, you should turn them into proper missing values so pandas ignores them in calculations. The code first makes a .copy() so the raw data stays untouched, then uses .replace(0, np.nan) on the five suspect columns only. Pregnancies and Outcome are left alone because zero is a valid value there. .isna().sum() then counts missing values, and the output matches the zero counts exactly (for example 374 for Insulin), confirming that every zero was converted and nothing else changed.
diabetes_clean = diabetes.copy()
diabetes_clean[check_cols] = diabetes_clean[check_cols].replace(0, np.nan)
diabetes_clean[check_cols].isna().sum()
Part 6: Real Descriptive Statistics#
With the zeros fixed, summary statistics finally describe real patients. describe() gives the count, mean, standard deviation, minimum, quartiles and maximum for every numeric column, and .round(2) tidies the decimals. Look at the count row first: Insulin has only 394 values and SkinThickness 541, so those columns are far less complete. The minimums are now plausible, for example Glucose starts at 44 and BMI at 18.2. The median age is 29, so this is a fairly young group of patients.
diabetes_clean.describe().round(2)
Part 7: Real Glucose by Diagnosis#
Comparing average glucose between patients with and without diabetes is the most direct way to see whether glucose separates the two groups. groupby('Outcome') splits the data by diagnosis, ['Glucose'].mean() averages each group, and .round(1) tidies the result. Missing values are skipped automatically by mean(). The output shows 110.6 for patients without diabetes and 142.3 for those with it, a gap of about 32 units. That is a large, clear difference and an early hint that glucose will be the strongest signal.
glucose_by_outcome = diabetes_clean.groupby('Outcome')['Glucose'].mean().round(1)
glucose_by_outcome
Part 8: Real BMI by Diagnosis#
Body weight is a well-known diabetes risk factor, so it is worth checking whether the data agrees. The code uses the same groupby('Outcome') pattern on BMI (body mass index, weight in kilograms divided by height in metres squared). The output shows an average BMI of 30.9 for patients without diabetes and 35.4 for those with it. Both groups are on average above 30, the usual cut-off for obesity, but the diagnosed group is clearly heavier. Reusing one pattern across columns like this keeps your analysis consistent and easy to check.
bmi_by_outcome = diabetes_clean.groupby('Outcome')['BMI'].mean().round(1)
bmi_by_outcome
Part 9: Real Age by Diagnosis#
Age matters because diabetes risk generally rises over time, and an age difference between groups can also explain other differences you see. The code averages Age within each Outcome group. The output shows 31.2 years for patients without diabetes and 37.1 for those with it, roughly six years apart. Keep this in mind for the pregnancies comparison that follows: older patients have had more time to have children, so the two measurements are likely to move together.
age_by_outcome = diabetes_clean.groupby('Outcome')['Age'].mean().round(1)
age_by_outcome
Part 10: Real Pregnancies by Diagnosis#
Pregnancy history is included in this dataset because pregnancy can affect how the body handles blood sugar. The code averages Pregnancies by Outcome, and the output shows 3.3 for patients without diabetes and 4.9 for those with it. Be careful how you read this. Patients with diabetes are also older on average, so part of this gap may simply reflect age. A difference in means shows association, not cause.
preg_by_outcome = diabetes_clean.groupby('Outcome')['Pregnancies'].mean().round(1)
preg_by_outcome
Part 11: Visualizing Real Glucose Distributions#
Averages hide the shape of the data. A histogram shows whether the two groups overlap heavily or sit apart. The code draws two plt.hist calls on the same axes, one per diagnosis, each with .dropna() to remove missing glucose values and bins=20. alpha=0.6 makes the bars semi-transparent so the overlap is visible. The chart is saved with plt.savefig and closed, so there is no inline output. When you open the image, look for the red diabetes bars sitting further right, and notice how much the two shapes still overlap in the middle range.
plt.figure(figsize=(8, 5))
plt.hist(diabetes_clean[diabetes_clean['Outcome']==0]['Glucose'].dropna(), bins=20, alpha=0.6, label='No Diabetes', color='seagreen')
plt.hist(diabetes_clean[diabetes_clean['Outcome']==1]['Glucose'].dropna(), bins=20, alpha=0.6, label='Diabetes', color='crimson')
plt.xlabel('Real Glucose Level')
plt.ylabel('Real Patient Count')
plt.title('Real Glucose Distribution by Diagnosis')
plt.legend()
plt.tight_layout()
plt.savefig('glucose_by_outcome.png', dpi=120)
plt.close()
Part 12: Real Correlation with Diagnosis#
Correlation gives you one number per column to rank how strongly each measurement moves with the diagnosis. corr(numeric_only=True) computes Pearson correlations (from -1 to 1, where 0 means no straight-line relationship) between all numeric columns, then ['Outcome'] keeps only the column against the target and sort_values(ascending=False) ranks it. The output puts Glucose well ahead at 0.495, followed by BMI at 0.314 and Insulin at 0.303, with BloodPressure lowest at 0.171. Ignore the 1.000 for Outcome itself: every column correlates perfectly with itself.
outcome_corr = diabetes_clean.corr(numeric_only=True)['Outcome'].sort_values(ascending=False)
outcome_corr.round(3)
Part 13: Visualizing the Real Correlation Matrix#
A heatmap of the full correlation matrix shows not only links to the diagnosis but also links between the measurements themselves. The code builds the matrix with corr(numeric_only=True), draws it with plt.imshow using the coolwarm colour map, and fixes the scale with vmin=-1, vmax=1 so colours mean the same thing in every chart. The chart is saved to a file, so there is no inline output. In the image, look for warmer squares away from the diagonal, such as age with pregnancies, which tell you two measurements carry overlapping information.
corr_matrix = diabetes_clean.corr(numeric_only=True)
plt.figure(figsize=(8, 7))
plt.imshow(corr_matrix, cmap='coolwarm', vmin=-1, vmax=1)
plt.colorbar(label='Real Correlation')
plt.xticks(range(len(corr_matrix.columns)), corr_matrix.columns, rotation=90)
plt.yticks(range(len(corr_matrix.columns)), corr_matrix.columns)
plt.title('Real Clinical Feature Correlation Matrix')
plt.tight_layout()
plt.savefig('correlation_heatmap.png', dpi=120)
plt.close()
Part 14: Real BMI Categories#
Clinicians talk about BMI in standard bands, not raw numbers, so grouping BMI makes results easier to explain. pd.cut assigns each value to a bin using the edges 0, 18.5, 25, 30 and 100 with the labels Underweight, Normal, Overweight and Obese. By default the right edge is included, so a BMI of exactly 30 counts as Overweight. value_counts() shows 465 obese patients, 180 overweight, 108 normal and only 4 underweight. Patients with missing BMI get no category and are not counted.
bmi_bins = [0, 18.5, 25, 30, 100]
bmi_labels = ['Underweight', 'Normal', 'Overweight', 'Obese']
diabetes_clean['bmi_category'] = pd.cut(diabetes_clean['BMI'], bins=bmi_bins, labels=bmi_labels)
diabetes_clean['bmi_category'].value_counts()
Part 15: Real Diabetes Rate by BMI Category#
Now you can see how the diagnosis rate changes from one BMI band to the next. The code groups by bmi_category and takes the mean of Outcome, which gives the share with diabetes, then multiplies by 100. observed=True tells pandas to show only categories that actually appear, which avoids a warning with categorical columns. The output climbs steadily: 0.0% for underweight, 6.5% for normal weight, 24.4% for overweight and 46.2% for obese. The underweight figure is based on only 4 patients, so do not read much into it.
bmi_diabetes_rate = (diabetes_clean.groupby('bmi_category', observed=True)['Outcome'].mean() * 100).round(1)
bmi_diabetes_rate
Part 16: Real Glucose Categories#
Glucose has its own clinical bands, so the same banding approach applies. pd.cut uses edges 0, 100, 126 and 300 with the labels Normal, Prediabetes and Diabetes Range, roughly mirroring common fasting glucose thresholds. The code then computes the diagnosis rate per band in one step. The output shows 8.6% for the normal band, 27.8% for prediabetes and 60.4% for the diabetes range. Each step up the scale multiplies the rate, which is a much easier story to tell a non-technical audience than a correlation of 0.495.
glucose_bins = [0, 100, 126, 300]
glucose_labels = ['Normal', 'Prediabetes', 'Diabetes Range']
diabetes_clean['glucose_category'] = pd.cut(diabetes_clean['Glucose'], bins=glucose_bins, labels=glucose_labels)
glucose_diabetes_rate = (diabetes_clean.groupby('glucose_category', observed=True)['Outcome'].mean() * 100).round(1)
glucose_diabetes_rate
Part 17: Real Age Groups#
Age groups show whether risk rises steadily with age or follows a different pattern. pd.cut uses edges 20, 30, 40, 50, 60 and 100, and right=False makes each bin include its left edge and exclude its right, so 30 falls in 30-39 rather than 21-29. The output rises from 21.2% for 21-29 to 46.1% for 30-39, 55.1% for 40-49 and 59.6% for 50-59, then drops to 28.1% for 60+. That last drop is surprising and is worth checking against how many patients are in the oldest group before drawing conclusions.
age_bins = [20, 30, 40, 50, 60, 100]
age_labels = ['21-29', '30-39', '40-49', '50-59', '60+']
diabetes_clean['age_group'] = pd.cut(diabetes_clean['Age'], bins=age_bins, labels=age_labels, right=False)
age_diabetes_rate = (diabetes_clean.groupby('age_group', observed=True)['Outcome'].mean() * 100).round(1)
age_diabetes_rate
Part 18: Visualizing Real Risk by Age Group#
A bar chart makes the age pattern obvious at a glance, which is what you would show a stakeholder. The code calls .plot(kind='bar') directly on the age rate Series, labels the axis, and uses plt.xticks(rotation=0) so the age labels sit horizontally. The figure is saved and closed, so there is no inline output. In the image, look for bars that climb steadily up to the 50-59 group and then fall sharply for 60+, the same pattern as the numbers in the previous step.
plt.figure(figsize=(8, 5))
age_diabetes_rate.plot(kind='bar', color='darkorange')
plt.ylabel('Real Diabetes Rate (%)')
plt.title('Real Diabetes Rate by Real Age Group')
plt.xticks(rotation=0)
plt.tight_layout()
plt.savefig('diabetes_rate_by_age.png', dpi=120)
plt.close()
Part 19: Real High-Risk Combination#
Risk factors rarely act alone, so it is useful to test what happens when two of them occur together. The code filters for patients who are both in the diabetes glucose range and in the obese BMI band, combining two conditions with & and wrapping each in brackets. It then returns the group size and its diagnosis rate. The output shows 209 patients with a rate of 70.8%, higher than the 60.4% for the glucose band alone and the 46.2% for obesity alone, and about double the overall rate of 34.9%.
high_risk = diabetes_clean[(diabetes_clean['glucose_category']=='Diabetes Range') & (diabetes_clean['bmi_category']=='Obese')]
high_risk_rate = round(high_risk['Outcome'].mean() * 100, 1)
len(high_risk), high_risk_rate
Part 20: Real Diabetes Pedigree Function#
The diabetes pedigree function is a score that summarises how much diabetes appears in a patient's family history, with higher values meaning a stronger family link. The code averages it within each Outcome group. The output shows 0.43 for patients without diabetes and 0.55 for those with it. The gap is real but modest, which fits its low correlation of 0.174 from earlier. Family history adds some information, but much less than glucose.
dpf_by_outcome = diabetes_clean.groupby('Outcome')['DiabetesPedigreeFunction'].mean().round(3)
dpf_by_outcome
Part 21: Real Skin Thickness by Diagnosis#
Skin thickness here is a skinfold measurement, often used as a rough indicator of body fat. The code averages SkinThickness by diagnosis, and because the zeros were converted to missing values earlier, only the 541 recorded measurements are used. The output shows 27.2 for patients without diabetes and 33.0 for those with it. Had you kept the 227 zeros, both averages would have been pulled down, which is exactly why the cleaning step mattered.
skin_by_outcome = diabetes_clean.groupby('Outcome')['SkinThickness'].mean().round(1)
skin_by_outcome
Part 22: Real Blood Pressure by Diagnosis#
Blood pressure completes the set of group comparisons. The code averages BloodPressure within each Outcome group. The output shows 70.9 for patients without diabetes and 75.3 for those with it, a gap of under 5 units. Compared with the glucose gap of over 30 units, this is small, which matches blood pressure having the weakest correlation with the diagnosis at 0.171. Not every known risk factor shows up strongly in every dataset.
bp_by_outcome = diabetes_clean.groupby('Outcome')['BloodPressure'].mean().round(1)
bp_by_outcome
Part 23: Real Simple Threshold Risk Rule#
Before building any model, it helps to know how well a single simple rule performs. The code flags a patient as high risk if glucose is above 140, using .fillna(0) so missing glucose values are treated as not flagged, and .astype(int) to turn True and False into 1 and 0. It then counts three outcomes by combining conditions with &. The output shows 132 true positives (flagged and diabetic), 60 false positives (flagged but not diabetic) and 136 false negatives (diabetic but missed).
risk_flag = (diabetes_clean['Glucose'].fillna(0) > 140).astype(int)
actual = diabetes_clean['Outcome']
true_positive = ((risk_flag == 1) & (actual == 1)).sum()
false_positive = ((risk_flag == 1) & (actual == 0)).sum()
false_negative = ((risk_flag == 0) & (actual == 1)).sum()
true_positive, false_positive, false_negative
Part 24: Real Precision and Recall#
Precision and recall turn those counts into two easy-to-read scores. Precision is the share of flagged patients who really have diabetes: true positives divided by everyone flagged. Recall is the share of diabetic patients the rule catches: true positives divided by all diabetic patients. The output gives a precision of 0.688 and a recall of 0.493. So about 69% of flagged patients are correctly flagged, but the rule misses roughly half of those with diabetes. In a screening setting, missing cases is often the more costly mistake.
precision = round(true_positive / (true_positive + false_positive), 3)
recall = round(true_positive / (true_positive + false_negative), 3)
precision, recall
Part 25: Real Missing Data Impact Check#
Missing data limits what you can trust, so it is worth putting a single number on the worst column. isna() marks each missing value as True, and .mean() on those booleans gives the share that is missing; multiplying by 100 makes it a percentage. The output shows 48.7% of Insulin values are missing. With nearly half the column absent, any conclusion based on insulin, including its 0.303 correlation, rests on a much smaller and possibly unrepresentative group of patients.
insulin_missing_pct = round(diabetes_clean['Insulin'].isna().mean() * 100, 1)
insulin_missing_pct
Part 26: Saving the Real Cleaned Dataset#
Saving the cleaned data means later work can start from the corrected version instead of repeating the zero fix. to_csv writes the file, and index=False stops pandas adding an extra unnamed column for the row index. The code then reads the file back and compares shapes. The output is True, so the saved file has the same number of rows and columns as the cleaned DataFrame. A quick reload test like this catches problems such as a forgotten index=False early.
diabetes_clean.to_csv('diabetes_cleaned.csv', index=False)
reloaded = pd.read_csv('diabetes_cleaned.csv')
reloaded.shape == diabetes_clean.shape
Part 27: Real Sanity Check, Outcome Rate#
Sanity checks confirm that a key number is what you expect, so a later edit cannot silently change it. The code compares the stored outcome_rate against 34.9 with ==, and the output is True. Comparing rounded decimals with == works here because the value was rounded to one decimal place earlier. With unrounded figures, use a tolerance such as np.isclose instead, since tiny floating-point differences can make an exact comparison fail.
outcome_rate == 34.9
Part 28: Real Sanity Check, Cleaning Did Not Drop Rows#
Cleaning steps can sometimes remove rows by accident, for example if you use dropna() without thinking. This check compares the number of rows in the cleaned data with the raw data using len(). The output is True, so all 768 patients are still present. That is the point of replacing zeros with missing values rather than deleting rows: you keep every patient and simply mark the unknown measurements.
len(diabetes_clean) == len(diabetes)
Part 29: Real Top Risk Factor Recommendation#
An analysis should end with a clear finding that someone can act on. The code drops the Outcome row from the correlation results, then uses .idxmax() to get the name of the highest remaining value and .max() for the value itself. An f-string prints the sentence. The output names Glucose as the strongest single correlate at 0.495. Building the message from variables rather than typing the numbers means it stays correct if the data changes.
top_factor = outcome_corr.drop('Outcome').idxmax()
top_factor_corr = round(outcome_corr.drop('Outcome').max(), 3)
print(f'{top_factor} is the real strongest single correlate of a real diabetes diagnosis in this dataset, at a real correlation of {top_factor_corr}.')
Part 30: Real Recap Print#
A one-sentence recap is useful for reports, emails and slide notes. The code prints an f-string combining the number of patients, the overall diagnosis rate and the top factor, all pulled from variables calculated earlier. The output reads: across 768 patients, 34.9% were diagnosed with diabetes, and glucose stood out as the strongest individual risk signal. Keep recaps like this short and factual, and let the detailed tables support them.
print(f'Across {len(diabetes_clean)} real patients, {outcome_rate}% were diagnosed with diabetes, and {top_factor} stood out as the real strongest individual risk signal.')
Wrap-Up: What You Learned#
- A real zero is not always a real zero, this dataset hid its real missing values inside a field that looks perfectly valid at first glance.
- Real glucose was the real single strongest correlate of a real diabetes diagnosis, well ahead of BMI, age, or pregnancy count.
- Converting real continuous fields like BMI, glucose, and age into real standard clinical categories made real risk patterns far easier to see and explain.
- Combining two real risk factors compounded real risk well beyond either factor alone.
- Even a real single-threshold rule, with no modeling at all, can be honestly scored with real precision and real recall.
- Next video: a real heart disease risk factor analysis, using a real classic clinical dataset from the Cleveland Clinic.
Practice on your own
- Repeat the banding approach for
Pregnancies: usepd.cutto create groups such as 0, 1-3, 4-6 and 7+, then calculate the diabetes rate per group withgroupbyandobserved=True. - Try different glucose thresholds in the simple risk rule, for example 120, 130 and 150, and record precision and recall for each. Note how raising the threshold trades recall for precision.
- Count how many patients fall in each age group with
value_counts()and check whether the drop in diabetes rate for the 60+ group is based on a small number of patients.
The next video moves on to a heart disease risk factor analysis using a classic clinical dataset from the Cleveland Clinic, where you can reuse the same cleaning and comparison steps.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



