Lesson 2 · Pandas Projects
Track Fitness & Health Data with Python Pandas | Hands-On Project
A complete, standalone tutorial: build six months of synthetic fitness tracker data, then explore trends, streaks, and correlations with pandas. No prior…
- CoursePandas Projects
- Lesson2 of 10
- Video18 min
- FormatJupyter notebook · 16 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbPandas for Health and Fitness: Analyzing a Daily Activity Log#
- A complete, standalone tutorial: build six months of synthetic fitness tracker data, then explore trends, streaks, and correlations with pandas.
- No prior pandas experience needed. Let's jump straight in.
Before You Start#
- Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
- If pandas isn't installed yet, open a terminal in VS Code and run: pip install pandas
Part 1: Building the Dataset#
import pandas as pd
import numpy as np
print(pd.__version__)
Generating a Synthetic Activity Log#
rng = np.random.default_rng(seed=77)
dates = pd.date_range('2026-02-01', periods=180, freq='D')
is_weekend = dates.dayofweek >= 5
active_minutes = np.where(is_weekend, rng.normal(70, 25, size=180), rng.normal(45, 18, size=180))
active_minutes = np.clip(active_minutes, 5, None).round(0)
print(active_minutes[:5])
steps = (active_minutes * rng.uniform(90, 130, size=180) + rng.normal(0, 800, size=180)).round(0).astype(int)
steps = np.clip(steps, 500, None)
calories_burned = (1600 + steps * 0.04 + active_minutes * 3 + rng.normal(0, 80, size=180)).round(0)
print(steps[:5])
sleep_hours = np.clip(rng.normal(7.0, 1.0, size=180), 3.5, 10.5).round(1)
resting_heart_rate = (68 - (sleep_hours - 7) * 2 + rng.normal(0, 3, size=180)).round(0).astype(int)
activity = pd.DataFrame({
'date': dates,
'steps': steps,
'active_minutes': active_minutes,
'calories_burned': calories_burned,
'sleep_hours': sleep_hours,
'resting_heart_rate': resting_heart_rate
})
print(activity.shape)
activity.to_csv('fitness_log.csv', index=False)
activity = pd.read_csv('fitness_log.csv', parse_dates=['date'])
activity.head()
Part 2: First Look at the Data#
print(activity.shape)
activity.info()
activity[['steps', 'active_minutes', 'calories_burned', 'sleep_hours', 'resting_heart_rate']].describe().round(1)
Part 3: Working with Dates#
activity['weekday'] = activity['date'].dt.day_name()
activity['is_weekend'] = activity['date'].dt.dayofweek >= 5
activity[['date', 'weekday', 'is_weekend']].head(3)
Part 4: Weekly Summaries and Rolling Averages#
weekly = activity.set_index('date').resample('W').agg(
total_steps=('steps', 'sum'),
avg_sleep=('sleep_hours', 'mean'),
avg_resting_hr=('resting_heart_rate', 'mean')
).round(1)
weekly.head()
activity['steps_rolling_7d'] = activity['steps'].rolling(window=7).mean().round(0)
activity[['date', 'steps', 'steps_rolling_7d']].tail()
Part 5: Goal Tracking and Streaks#
activity['hit_goal'] = activity['steps'] >= 8000
print(f"Days goal was hit: {activity['hit_goal'].sum()} out of {len(activity)}")
print(f"Goal hit rate: {activity['hit_goal'].mean() * 100:.1f}%")
streak_id = (activity['hit_goal'] != activity['hit_goal'].shift()).cumsum()
streak_lengths = activity[activity['hit_goal']].groupby(streak_id).size()
print(f'Longest goal streak: {streak_lengths.max()} days')
Part 6: Correlation and Weekday Comparison#
numeric_cols = ['steps', 'active_minutes', 'calories_burned', 'sleep_hours', 'resting_heart_rate']
activity[numeric_cols].corr().round(2)
weekday_comparison = activity.groupby('is_weekend').agg(
avg_steps=('steps', 'mean'),
avg_active_minutes=('active_minutes', 'mean'),
avg_sleep=('sleep_hours', 'mean')
).round(1)
weekday_comparison.index = ['Weekday', 'Weekend']
weekday_comparison
Part 7: Visualizing the Results#
import matplotlib
matplotlib.use('Agg')
import matplotlib.pyplot as plt
fig, ax = plt.subplots(figsize=(10, 5))
ax.plot(activity['date'], activity['steps'], color='gray', alpha=0.4, label='Daily steps')
ax.plot(activity['date'], activity['steps_rolling_7d'], color='seagreen', linewidth=2, label='7-day rolling average')
ax.axhline(8000, color='crimson', linestyle='--', linewidth=1, label='Goal (8,000)')
ax.set_title('Daily Steps Over Six Months')
ax.set_ylabel('Steps')
ax.legend()
plt.tight_layout()
plt.savefig('steps_trend.png', dpi=150)
plt.close(fig)
print('Saved steps_trend.png')
fig, ax = plt.subplots(figsize=(8, 5))
ax.scatter(activity['sleep_hours'], activity['resting_heart_rate'], alpha=0.5, color='mediumpurple')
ax.set_title('Sleep Hours vs. Resting Heart Rate')
ax.set_xlabel('Sleep Hours')
ax.set_ylabel('Resting Heart Rate (bpm)')
plt.tight_layout()
plt.savefig('sleep_vs_heart_rate.png', dpi=150)
plt.close(fig)
print('Saved sleep_vs_heart_rate.png')
Wrap-Up: What You Learned#
- Generating a realistic synthetic activity log with genuinely correlated metrics, then saving and reloading with to_csv and read_csv.
- First-look exploration with info and describe.
- Extracting weekday names and a weekend flag with the dt accessor.
- Weekly summaries with resample plus agg, and smoothing noisy days with a rolling average.
- Goal tracking with a boolean flag, plus a shift-and-cumsum trick for measuring real streaks.
- corr for uncovering relationships between metrics, and a weekday-versus-weekend comparison.
- You went from a raw synthetic activity feed to a full fitness analysis with streaks, correlations, and charts. If you want the next dataset in this series to land in your feed automatically, subscribing is the move see you in the next one.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



