Mathew K Analytics

Lesson 2 · Pandas Projects

Track Fitness & Health Data with Python Pandas | Hands-On Project

A complete, standalone tutorial: build six months of synthetic fitness tracker data, then explore trends, streaks, and correlations with pandas. No prior…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Pandas for Health and Fitness: Analyzing a Daily Activity Log#

  • A complete, standalone tutorial: build six months of synthetic fitness tracker data, then explore trends, streaks, and correlations with pandas.
  • No prior pandas experience needed. Let's jump straight in.

Before You Start#

  • Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
  • If pandas isn't installed yet, open a terminal in VS Code and run: pip install pandas

Part 1: Building the Dataset#

import pandas as pd
import numpy as np
print(pd.__version__)
2.3.0

Generating a Synthetic Activity Log#

rng = np.random.default_rng(seed=77)
dates = pd.date_range('2026-02-01', periods=180, freq='D')
is_weekend = dates.dayofweek >= 5

active_minutes = np.where(is_weekend, rng.normal(70, 25, size=180), rng.normal(45, 18, size=180))
active_minutes = np.clip(active_minutes, 5, None).round(0)
print(active_minutes[:5])
[81. 66. 31. 49. 44.]
steps = (active_minutes * rng.uniform(90, 130, size=180) + rng.normal(0, 800, size=180)).round(0).astype(int)
steps = np.clip(steps, 500, None)
calories_burned = (1600 + steps * 0.04 + active_minutes * 3 + rng.normal(0, 80, size=180)).round(0)
print(steps[:5])
[8455 9296 3488 5849 4055]
sleep_hours = np.clip(rng.normal(7.0, 1.0, size=180), 3.5, 10.5).round(1)
resting_heart_rate = (68 - (sleep_hours - 7) * 2 + rng.normal(0, 3, size=180)).round(0).astype(int)

activity = pd.DataFrame({
    'date': dates,
    'steps': steps,
    'active_minutes': active_minutes,
    'calories_burned': calories_burned,
    'sleep_hours': sleep_hours,
    'resting_heart_rate': resting_heart_rate
})
print(activity.shape)
(180, 6)
activity.to_csv('fitness_log.csv', index=False)
activity = pd.read_csv('fitness_log.csv', parse_dates=['date'])
activity.head()
date steps active_minutes calories_burned sleep_hours resting_heart_rate
0 2026-02-01 8455 81.0 2241.0 7.7 72
1 2026-02-02 9296 66.0 2067.0 5.9 69
2 2026-02-03 3488 31.0 1852.0 7.6 67
3 2026-02-04 5849 49.0 1954.0 6.1 72
4 2026-02-05 4055 44.0 2016.0 6.6 71

Part 2: First Look at the Data#

print(activity.shape)
activity.info()
(180, 6)
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 180 entries, 0 to 179
Data columns (total 6 columns):
 #   Column              Non-Null Count  Dtype         
---  ------              --------------  -----         
 0   date                180 non-null    datetime64[ns]
 1   steps               180 non-null    int64         
 2   active_minutes      180 non-null    float64       
 3   calories_burned     180 non-null    float64       
 4   sleep_hours         180 non-null    float64       
 5   resting_heart_rate  180 non-null    int64         
dtypes: datetime64[ns](1), float64(3), int64(2)
memory usage: 8.6 KB
activity[['steps', 'active_minutes', 'calories_burned', 'sleep_hours', 'resting_heart_rate']].describe().round(1)
steps active_minutes calories_burned sleep_hours resting_heart_rate
count 180.0 180.0 180.0 180.0 180.0
mean 5649.2 51.1 1979.5 7.0 68.0
std 2593.9 23.5 189.5 1.0 3.8
min 500.0 5.0 1482.0 4.2 55.0
25% 3921.2 35.0 1872.0 6.2 66.0
50% 5127.5 48.0 1959.5 7.1 68.0
75% 7479.2 65.0 2092.8 7.8 71.0
max 14917.0 162.0 2668.0 9.5 75.0

Part 3: Working with Dates#

activity['weekday'] = activity['date'].dt.day_name()
activity['is_weekend'] = activity['date'].dt.dayofweek >= 5
activity[['date', 'weekday', 'is_weekend']].head(3)
date weekday is_weekend
0 2026-02-01 Sunday True
1 2026-02-02 Monday False
2 2026-02-03 Tuesday False

Part 4: Weekly Summaries and Rolling Averages#

weekly = activity.set_index('date').resample('W').agg(
    total_steps=('steps', 'sum'),
    avg_sleep=('sleep_hours', 'mean'),
    avg_resting_hr=('resting_heart_rate', 'mean')
).round(1)
weekly.head()
total_steps avg_sleep avg_resting_hr
date
2026-02-01 8455 7.7 72.0
2026-02-08 46740 6.5 70.7
2026-02-15 37844 6.9 69.3
2026-02-22 34319 7.7 64.6
2026-03-01 38059 6.8 69.3
activity['steps_rolling_7d'] = activity['steps'].rolling(window=7).mean().round(0)
activity[['date', 'steps', 'steps_rolling_7d']].tail()
date steps steps_rolling_7d
175 2026-07-26 5241 5394.0
176 2026-07-27 7656 6042.0
177 2026-07-28 500 4760.0
178 2026-07-29 6160 4835.0
179 2026-07-30 4925 4990.0

Part 5: Goal Tracking and Streaks#

activity['hit_goal'] = activity['steps'] >= 8000
print(f"Days goal was hit: {activity['hit_goal'].sum()} out of {len(activity)}")
print(f"Goal hit rate: {activity['hit_goal'].mean() * 100:.1f}%")
Days goal was hit: 33 out of 180
Goal hit rate: 18.3%
streak_id = (activity['hit_goal'] != activity['hit_goal'].shift()).cumsum()
streak_lengths = activity[activity['hit_goal']].groupby(streak_id).size()
print(f'Longest goal streak: {streak_lengths.max()} days')
Longest goal streak: 3 days

Part 6: Correlation and Weekday Comparison#

numeric_cols = ['steps', 'active_minutes', 'calories_burned', 'sleep_hours', 'resting_heart_rate']
activity[numeric_cols].corr().round(2)
steps active_minutes calories_burned sleep_hours resting_heart_rate
steps 1.00 0.93 0.89 -0.02 -0.02
active_minutes 0.93 1.00 0.89 0.01 -0.05
calories_burned 0.89 0.89 1.00 -0.02 -0.05
sleep_hours -0.02 0.01 -0.02 1.00 -0.61
resting_heart_rate -0.02 -0.05 -0.05 -0.61 1.00
weekday_comparison = activity.groupby('is_weekend').agg(
    avg_steps=('steps', 'mean'),
    avg_active_minutes=('active_minutes', 'mean'),
    avg_sleep=('sleep_hours', 'mean')
).round(1)
weekday_comparison.index = ['Weekday', 'Weekend']
weekday_comparison
avg_steps avg_active_minutes avg_sleep
Weekday 4927.6 44.1 6.9
Weekend 7474.4 68.7 7.2

Part 7: Visualizing the Results#

import matplotlib
matplotlib.use('Agg')
import matplotlib.pyplot as plt

fig, ax = plt.subplots(figsize=(10, 5))
ax.plot(activity['date'], activity['steps'], color='gray', alpha=0.4, label='Daily steps')
ax.plot(activity['date'], activity['steps_rolling_7d'], color='seagreen', linewidth=2, label='7-day rolling average')
ax.axhline(8000, color='crimson', linestyle='--', linewidth=1, label='Goal (8,000)')
ax.set_title('Daily Steps Over Six Months')
ax.set_ylabel('Steps')
ax.legend()
plt.tight_layout()
plt.savefig('steps_trend.png', dpi=150)
plt.close(fig)
print('Saved steps_trend.png')
Saved steps_trend.png
fig, ax = plt.subplots(figsize=(8, 5))
ax.scatter(activity['sleep_hours'], activity['resting_heart_rate'], alpha=0.5, color='mediumpurple')
ax.set_title('Sleep Hours vs. Resting Heart Rate')
ax.set_xlabel('Sleep Hours')
ax.set_ylabel('Resting Heart Rate (bpm)')
plt.tight_layout()
plt.savefig('sleep_vs_heart_rate.png', dpi=150)
plt.close(fig)
print('Saved sleep_vs_heart_rate.png')
Saved sleep_vs_heart_rate.png

Wrap-Up: What You Learned#

  • Generating a realistic synthetic activity log with genuinely correlated metrics, then saving and reloading with to_csv and read_csv.
  • First-look exploration with info and describe.
  • Extracting weekday names and a weekend flag with the dt accessor.
  • Weekly summaries with resample plus agg, and smoothing noisy days with a rolling average.
  • Goal tracking with a boolean flag, plus a shift-and-cumsum trick for measuring real streaks.
  • corr for uncovering relationships between metrics, and a weekday-versus-weekend comparison.
  • You went from a raw synthetic activity feed to a full fitness analysis with streaks, correlations, and charts. If you want the next dataset in this series to land in your feed automatically, subscribing is the move see you in the next one.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.