Mathew K Analytics

Lesson 5 · Python for Data Analysts

Time Series Analysis with Pandas

Everything you need to work with dates and time-indexed data: DatetimeIndex, resampling, rolling windows, seasonality, and simple forecasting. No prior time…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Time Series Analysis with Pandas#

  • Everything you need to work with dates and time-indexed data: DatetimeIndex, resampling, rolling windows, seasonality, and simple forecasting.
  • No prior time series experience needed. Let's get straight into it.

Before You Start#

  • Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
  • If pandas isn't installed yet, open a terminal in VS Code and run: pip install pandas

Part 1: DatetimeIndex and date_range#

import pandas as pd
import numpy as np

dates = pd.date_range(start='2024-01-01', periods=10, freq='D')
print(dates)
DatetimeIndex(['2024-01-01', '2024-01-02', '2024-01-03', '2024-01-04',
               '2024-01-05', '2024-01-06', '2024-01-07', '2024-01-08',
               '2024-01-09', '2024-01-10'],
              dtype='datetime64[ns]', freq='D')
weekly = pd.date_range(start='2024-01-01', periods=6, freq='W')
print(weekly)
month_end = pd.date_range(start='2024-01-01', periods=4, freq='ME')
print(month_end)
DatetimeIndex(['2024-01-07', '2024-01-14', '2024-01-21', '2024-01-28',
               '2024-02-04', '2024-02-11'],
              dtype='datetime64[ns]', freq='W-SUN')
DatetimeIndex(['2024-01-31', '2024-02-29', '2024-03-31', '2024-04-30'], dtype='datetime64[ns]', freq='ME')
raw_dates = ['2024-01-01', '2024-02-15', '2024-03-30']
parsed = pd.to_datetime(raw_dates)
print(parsed)
print(parsed.dtype)
DatetimeIndex(['2024-01-01', '2024-02-15', '2024-03-30'], dtype='datetime64[ns]', freq=None)
datetime64[ns]
sample = pd.Series(range(10), index=dates)
print(sample.index.year)
print(sample.index.day_name())
Index([2024, 2024, 2024, 2024, 2024, 2024, 2024, 2024, 2024, 2024], dtype='int32')
Index(['Monday', 'Tuesday', 'Wednesday', 'Thursday', 'Friday', 'Saturday',
       'Sunday', 'Monday', 'Tuesday', 'Wednesday'],
      dtype='object')

Part 2: Indexing and Slicing Time Series#

long_dates = pd.date_range(start='2024-01-01', periods=90, freq='D')
rng = np.random.default_rng(seed=3)
ts = pd.Series(rng.integers(50, 150, size=90), index=long_dates)
print(ts.head())
2024-01-01    131
2024-01-02     58
2024-01-03     67
2024-01-04     73
2024-01-05     68
Freq: D, dtype: int64
print(ts['2024-01-05'])
print(ts['2024-02'].head())
68
2024-02-01    145
2024-02-02    128
2024-02-03     78
2024-02-04     81
2024-02-05    114
Freq: D, dtype: int64
january = ts['2024-01-01':'2024-01-31']
print(len(january))
print(january.loc['2024-01-15':'2024-01-20'])
31
2024-01-15     76
2024-01-16     65
2024-01-17    119
2024-01-18    123
2024-01-19     53
2024-01-20     61
Freq: D, dtype: int64

Part 3: Resampling#

weekly_totals = ts.resample('W').sum()
print(weekly_totals.head())
2024-01-07    663
2024-01-14    605
2024-01-21    592
2024-01-28    737
2024-02-04    747
Freq: W-SUN, dtype: int64
monthly_stats = ts.resample('ME').agg(['mean', 'min', 'max'])
print(monthly_stats)
                  mean  min  max
2024-01-31   93.935484   53  138
2024-02-29   99.586207   50  147
2024-03-31  102.133333   50  143
sparse_dates = pd.to_datetime(['2024-01-01', '2024-01-04', '2024-01-08'])
sparse = pd.Series([100, 130, 90], index=sparse_dates)
daily_filled = sparse.resample('D').asfreq()
print(daily_filled)
2024-01-01    100.0
2024-01-02      NaN
2024-01-03      NaN
2024-01-04    130.0
2024-01-05      NaN
2024-01-06      NaN
2024-01-07      NaN
2024-01-08     90.0
Freq: D, dtype: float64
daily_forward_filled = sparse.resample('D').ffill()
print(daily_forward_filled)
2024-01-01    100
2024-01-02    100
2024-01-03    100
2024-01-04    130
2024-01-05    130
2024-01-06    130
2024-01-07    130
2024-01-08     90
Freq: D, dtype: int64

Part 4: Rolling and Expanding Windows#

rolling_avg = ts.rolling(window=7).mean()
print(rolling_avg.head(10))
2024-01-01          NaN
2024-01-02          NaN
2024-01-03          NaN
2024-01-04          NaN
2024-01-05          NaN
2024-01-06          NaN
2024-01-07    94.714286
2024-01-08    91.428571
2024-01-09    90.714286
2024-01-10    89.571429
Freq: D, dtype: float64
rolling_avg_partial = ts.rolling(window=7, min_periods=1).mean()
print(rolling_avg_partial.head(10))
2024-01-01    131.000000
2024-01-02     94.500000
2024-01-03     85.333333
2024-01-04     82.250000
2024-01-05     79.400000
2024-01-06     87.833333
2024-01-07     94.714286
2024-01-08     91.428571
2024-01-09     90.714286
2024-01-10     89.571429
Freq: D, dtype: float64
rolling_std = ts.rolling(window=7).std()
print(rolling_std.head(10))
2024-01-01          NaN
2024-01-02          NaN
2024-01-03          NaN
2024-01-04          NaN
2024-01-05          NaN
2024-01-06          NaN
2024-01-07    35.513914
2024-01-08    32.536426
2024-01-09    33.435083
2024-01-10    34.500518
Freq: D, dtype: float64
expanding_avg = ts.expanding().mean()
print(expanding_avg.head(10))
2024-01-01    131.000000
2024-01-02     94.500000
2024-01-03     85.333333
2024-01-04     82.250000
2024-01-05     79.400000
2024-01-06     87.833333
2024-01-07     94.714286
2024-01-08     96.375000
2024-01-09     91.555556
2024-01-10     88.300000
Freq: D, dtype: float64

Part 5: Shift, Diff, and Percent Change#

shifted = ts.shift(1)
comparison = pd.DataFrame({'today': ts, 'yesterday': shifted})
print(comparison.head())
            today  yesterday
2024-01-01    131        NaN
2024-01-02     58      131.0
2024-01-03     67       58.0
2024-01-04     73       67.0
2024-01-05     68       73.0
day_over_day = ts.diff()
print(day_over_day.head())
2024-01-01     NaN
2024-01-02   -73.0
2024-01-03     9.0
2024-01-04     6.0
2024-01-05    -5.0
Freq: D, dtype: float64
pct = ts.pct_change() * 100
print(pct.head())
2024-01-01          NaN
2024-01-02   -55.725191
2024-01-03    15.517241
2024-01-04     8.955224
2024-01-05    -6.849315
Freq: D, dtype: float64
week_over_week = ts.diff(periods=7)
print(week_over_week.head(10))
2024-01-01     NaN
2024-01-02     NaN
2024-01-03     NaN
2024-01-04     NaN
2024-01-05     NaN
2024-01-06     NaN
2024-01-07     NaN
2024-01-08   -23.0
2024-01-09    -5.0
2024-01-10    -8.0
Freq: D, dtype: float64

Part 6: Timezones#

naive = pd.Timestamp('2024-06-15 09:00:00')
print(naive)
print(naive.tz)
2024-06-15 09:00:00
None
aware = naive.tz_localize('America/New_York')
print(aware)
print(aware.tz)
2024-06-15 09:00:00-04:00
America/New_York
london_time = aware.tz_convert('Europe/London')
print(london_time)
tokyo_time = aware.tz_convert('Asia/Tokyo')
print(tokyo_time)
2024-06-15 14:00:00+01:00
2024-06-15 22:00:00+09:00

Part 7: Period vs. Timestamp#

point = pd.Timestamp('2024-03-15')
span = pd.Period('2024-03', freq='M')
print(point)
print(span)
print(span.start_time, '->', span.end_time)
2024-03-15 00:00:00
2024-03
2024-03-01 00:00:00 -> 2024-03-31 23:59:59.999999999
next_month = span + 1
print(next_month)
quarterly = pd.Period('2024Q2', freq='Q')
print(quarterly)
2024-04
2024Q2
monthly_series = ts.resample('ME').sum()
as_periods = monthly_series.to_period('M')
print(as_periods.index)
back_to_timestamps = as_periods.to_timestamp()
print(back_to_timestamps.index)
PeriodIndex(['2024-01', '2024-02', '2024-03'], dtype='period[M]')
DatetimeIndex(['2024-01-01', '2024-02-01', '2024-03-01'], dtype='datetime64[ns]', freq='MS')

Part 8: Finding Seasonal Patterns#

by_weekday = ts.groupby(ts.index.day_name()).mean()
weekday_order = ['Monday', 'Tuesday', 'Wednesday', 'Thursday', 'Friday', 'Saturday', 'Sunday']
print(by_weekday.reindex(weekday_order))
Monday        93.769231
Tuesday      101.846154
Wednesday     96.153846
Thursday     102.000000
Friday        99.153846
Saturday     105.307692
Sunday        90.583333
dtype: float64
by_month = ts.groupby(ts.index.month).mean()
print(by_month)
1     93.935484
2     99.586207
3    102.133333
dtype: float64

Part 9: Simple Forecasting#

naive_forecast = ts.iloc[-1]
print(f'Naive forecast for every future day: {naive_forecast}')
Naive forecast for every future day: 92
moving_avg_forecast = ts.iloc[-7:].mean()
print(f'7-day moving average forecast: {moving_avg_forecast:.2f}')
7-day moving average forecast: 102.86
days_numeric = np.arange(len(ts))
slope, intercept = np.polyfit(days_numeric, ts.values, deg=1)
print(f'Trend: {slope:.3f} units per day, starting around {intercept:.1f}')
Trend: 0.133 units per day, starting around 92.6
future_days = np.arange(len(ts), len(ts) + 14)
trend_forecast = slope * future_days + intercept
future_dates = pd.date_range(start=ts.index[-1] + pd.Timedelta(days=1), periods=14, freq='D')
forecast_series = pd.Series(trend_forecast, index=future_dates)
print(forecast_series)
2024-03-31    104.560799
2024-04-01    104.694248
2024-04-02    104.827696
2024-04-03    104.961145
2024-04-04    105.094593
2024-04-05    105.228042
2024-04-06    105.361490
2024-04-07    105.494939
2024-04-08    105.628388
2024-04-09    105.761836
2024-04-10    105.895285
2024-04-11    106.028733
2024-04-12    106.162182
2024-04-13    106.295630
Freq: D, dtype: float64

Capstone Project: Sales Trend Analysis and Forecast#

rng = np.random.default_rng(seed=21)
two_years = pd.date_range(start='2023-01-01', periods=730, freq='D')
day_index = np.arange(730)

trend = 200 + day_index * 0.15
weekday_effect = np.where(pd.Series(two_years).dt.dayofweek.values >= 5, 60, 0)
noise = rng.normal(0, 15, size=730)

daily_sales = pd.Series(trend + weekday_effect + noise, index=two_years).round(2)
print(daily_sales.head())
print(len(daily_sales), 'days generated')
2023-01-01    265.38
2023-01-02    222.81
2023-01-03    173.51
2023-01-04    225.75
2023-01-05    199.89
Freq: D, dtype: float64
730 days generated
def analyze_sales_trend(series, forecast_days=14):
    weekly_totals = series.resample('W').sum()
    rolling_4week_avg = weekly_totals.rolling(window=4, min_periods=1).mean()

    overall_avg = series.mean()
    weekday_avg = series.groupby(series.index.day_name()).mean()
    weekday_factor = weekday_avg / overall_avg

    recent_avg = series.iloc[-28:].mean()
    last_date = series.index[-1]
    future_dates = pd.date_range(start=last_date + pd.Timedelta(days=1), periods=forecast_days, freq='D')

    forecast_values = []
    for future_date in future_dates:
        factor = weekday_factor[future_date.day_name()]
        forecast_values.append(recent_avg * factor)
    forecast = pd.Series(forecast_values, index=future_dates).round(2)

    return {
        'weekly_totals': weekly_totals,
        'rolling_4week_avg': rolling_4week_avg,
        'weekday_factor': weekday_factor,
        'forecast': forecast,
    }
results = analyze_sales_trend(daily_sales, forecast_days=14)
print(results['weekday_factor'])
Friday       0.935465
Monday       0.945705
Saturday     1.158459
Sunday       1.154706
Thursday     0.933546
Tuesday      0.930189
Wednesday    0.940965
dtype: float64
print(results['rolling_4week_avg'].tail())
print()
print(results['forecast'])
2024-12-08    2259.7175
2024-12-15    2288.8350
2024-12-22    2296.7550
2024-12-29    2308.1375
2025-01-05    1814.4675
Freq: W-SUN, dtype: float64

2024-12-31    307.36
2025-01-01    310.92
2025-01-02    308.47
2025-01-03    309.10
2025-01-04    382.78
2025-01-05    381.54
2025-01-06    312.48
2025-01-07    307.36
2025-01-08    310.92
2025-01-09    308.47
2025-01-10    309.10
2025-01-11    382.78
2025-01-12    381.54
2025-01-13    312.48
Freq: D, dtype: float64

Wrap-Up: What You Learned#

  • DatetimeIndex and date_range, plus parsing dates with to_datetime.
  • Indexing and slicing time series with partial string indexing and date ranges.
  • Resampling: downsampling with aggregations, and upsampling with asfreq and ffill.
  • Rolling and expanding windows for smoothing and cumulative views.
  • shift, diff, and pct_change, including period-over-period comparisons.
  • Timezones: localizing and converting between them.
  • Period versus Timestamp, for representing spans of time versus exact instants.
  • Finding seasonal patterns by grouping on weekday or month.
  • Simple forecasting baselines: naive, moving average, and linear trend.
  • A capstone pipeline combining trend, seasonality, and forecasting into one reusable function.
  • You went from a plain date string to a seasonality-aware sales forecast in one sitting. If you want the next build to land in your feed automatically, subscribing is the move see you in the next one.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.