Lesson 5 · Python for Data Analysts
Time Series Analysis with Pandas
Everything you need to work with dates and time-indexed data: DatetimeIndex, resampling, rolling windows, seasonality, and simple forecasting. No prior time…
- CoursePython for Data Analysts
- Lesson5 of 12
- Video34 min
- FormatJupyter notebook · 35 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbTime Series Analysis with Pandas#
- Everything you need to work with dates and time-indexed data: DatetimeIndex, resampling, rolling windows, seasonality, and simple forecasting.
- No prior time series experience needed. Let's get straight into it.
Before You Start#
- Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
- If pandas isn't installed yet, open a terminal in VS Code and run: pip install pandas
Part 1: DatetimeIndex and date_range#
import pandas as pd
import numpy as np
dates = pd.date_range(start='2024-01-01', periods=10, freq='D')
print(dates)
weekly = pd.date_range(start='2024-01-01', periods=6, freq='W')
print(weekly)
month_end = pd.date_range(start='2024-01-01', periods=4, freq='ME')
print(month_end)
raw_dates = ['2024-01-01', '2024-02-15', '2024-03-30']
parsed = pd.to_datetime(raw_dates)
print(parsed)
print(parsed.dtype)
sample = pd.Series(range(10), index=dates)
print(sample.index.year)
print(sample.index.day_name())
Part 2: Indexing and Slicing Time Series#
long_dates = pd.date_range(start='2024-01-01', periods=90, freq='D')
rng = np.random.default_rng(seed=3)
ts = pd.Series(rng.integers(50, 150, size=90), index=long_dates)
print(ts.head())
print(ts['2024-01-05'])
print(ts['2024-02'].head())
january = ts['2024-01-01':'2024-01-31']
print(len(january))
print(january.loc['2024-01-15':'2024-01-20'])
Part 3: Resampling#
weekly_totals = ts.resample('W').sum()
print(weekly_totals.head())
monthly_stats = ts.resample('ME').agg(['mean', 'min', 'max'])
print(monthly_stats)
sparse_dates = pd.to_datetime(['2024-01-01', '2024-01-04', '2024-01-08'])
sparse = pd.Series([100, 130, 90], index=sparse_dates)
daily_filled = sparse.resample('D').asfreq()
print(daily_filled)
daily_forward_filled = sparse.resample('D').ffill()
print(daily_forward_filled)
Part 4: Rolling and Expanding Windows#
rolling_avg = ts.rolling(window=7).mean()
print(rolling_avg.head(10))
rolling_avg_partial = ts.rolling(window=7, min_periods=1).mean()
print(rolling_avg_partial.head(10))
rolling_std = ts.rolling(window=7).std()
print(rolling_std.head(10))
expanding_avg = ts.expanding().mean()
print(expanding_avg.head(10))
Part 5: Shift, Diff, and Percent Change#
shifted = ts.shift(1)
comparison = pd.DataFrame({'today': ts, 'yesterday': shifted})
print(comparison.head())
day_over_day = ts.diff()
print(day_over_day.head())
pct = ts.pct_change() * 100
print(pct.head())
week_over_week = ts.diff(periods=7)
print(week_over_week.head(10))
Part 6: Timezones#
naive = pd.Timestamp('2024-06-15 09:00:00')
print(naive)
print(naive.tz)
aware = naive.tz_localize('America/New_York')
print(aware)
print(aware.tz)
london_time = aware.tz_convert('Europe/London')
print(london_time)
tokyo_time = aware.tz_convert('Asia/Tokyo')
print(tokyo_time)
Part 7: Period vs. Timestamp#
point = pd.Timestamp('2024-03-15')
span = pd.Period('2024-03', freq='M')
print(point)
print(span)
print(span.start_time, '->', span.end_time)
next_month = span + 1
print(next_month)
quarterly = pd.Period('2024Q2', freq='Q')
print(quarterly)
monthly_series = ts.resample('ME').sum()
as_periods = monthly_series.to_period('M')
print(as_periods.index)
back_to_timestamps = as_periods.to_timestamp()
print(back_to_timestamps.index)
Part 8: Finding Seasonal Patterns#
by_weekday = ts.groupby(ts.index.day_name()).mean()
weekday_order = ['Monday', 'Tuesday', 'Wednesday', 'Thursday', 'Friday', 'Saturday', 'Sunday']
print(by_weekday.reindex(weekday_order))
by_month = ts.groupby(ts.index.month).mean()
print(by_month)
Part 9: Simple Forecasting#
naive_forecast = ts.iloc[-1]
print(f'Naive forecast for every future day: {naive_forecast}')
moving_avg_forecast = ts.iloc[-7:].mean()
print(f'7-day moving average forecast: {moving_avg_forecast:.2f}')
days_numeric = np.arange(len(ts))
slope, intercept = np.polyfit(days_numeric, ts.values, deg=1)
print(f'Trend: {slope:.3f} units per day, starting around {intercept:.1f}')
future_days = np.arange(len(ts), len(ts) + 14)
trend_forecast = slope * future_days + intercept
future_dates = pd.date_range(start=ts.index[-1] + pd.Timedelta(days=1), periods=14, freq='D')
forecast_series = pd.Series(trend_forecast, index=future_dates)
print(forecast_series)
Capstone Project: Sales Trend Analysis and Forecast#
rng = np.random.default_rng(seed=21)
two_years = pd.date_range(start='2023-01-01', periods=730, freq='D')
day_index = np.arange(730)
trend = 200 + day_index * 0.15
weekday_effect = np.where(pd.Series(two_years).dt.dayofweek.values >= 5, 60, 0)
noise = rng.normal(0, 15, size=730)
daily_sales = pd.Series(trend + weekday_effect + noise, index=two_years).round(2)
print(daily_sales.head())
print(len(daily_sales), 'days generated')
def analyze_sales_trend(series, forecast_days=14):
weekly_totals = series.resample('W').sum()
rolling_4week_avg = weekly_totals.rolling(window=4, min_periods=1).mean()
overall_avg = series.mean()
weekday_avg = series.groupby(series.index.day_name()).mean()
weekday_factor = weekday_avg / overall_avg
recent_avg = series.iloc[-28:].mean()
last_date = series.index[-1]
future_dates = pd.date_range(start=last_date + pd.Timedelta(days=1), periods=forecast_days, freq='D')
forecast_values = []
for future_date in future_dates:
factor = weekday_factor[future_date.day_name()]
forecast_values.append(recent_avg * factor)
forecast = pd.Series(forecast_values, index=future_dates).round(2)
return {
'weekly_totals': weekly_totals,
'rolling_4week_avg': rolling_4week_avg,
'weekday_factor': weekday_factor,
'forecast': forecast,
}
results = analyze_sales_trend(daily_sales, forecast_days=14)
print(results['weekday_factor'])
print(results['rolling_4week_avg'].tail())
print()
print(results['forecast'])
Wrap-Up: What You Learned#
- DatetimeIndex and date_range, plus parsing dates with to_datetime.
- Indexing and slicing time series with partial string indexing and date ranges.
- Resampling: downsampling with aggregations, and upsampling with asfreq and ffill.
- Rolling and expanding windows for smoothing and cumulative views.
- shift, diff, and pct_change, including period-over-period comparisons.
- Timezones: localizing and converting between them.
- Period versus Timestamp, for representing spans of time versus exact instants.
- Finding seasonal patterns by grouping on weekday or month.
- Simple forecasting baselines: naive, moving average, and linear trend.
- A capstone pipeline combining trend, seasonality, and forecasting into one reusable function.
- You went from a plain date string to a seasonality-aware sales forecast in one sitting. If you want the next build to land in your feed automatically, subscribing is the move see you in the next one.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



