Mathew K Analytics

Lesson 36 · Mastering Pandas

Mastering Time Series Analysis in Python: Resampling, Shifting, and Rolling Windows with Pandas

Welcome! In this lesson, you will gain hands-on practice with intermediate time-series tasks in pandas. You will use the airline passenger Flights dataset…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Pandas Time-Series: Resampling, Shifting, and Rolling Windows#

Welcome! In this lesson, you will gain hands-on practice with intermediate time-series tasks in pandas.

You will use the airline passenger Flights dataset to learn about:

  • Resampling data to different time periods
  • Shifting values for comparisons
  • Calculating rolling window statistics

These are core skills for analyzing trends, smoothing, and detecting seasonality in real-world data.

import warnings; warnings.filterwarnings("ignore")  # Suppress warnings for clean output
import numpy as np
import pandas as pd
np.random.seed(42)
# Data setup (Airline Passenger Flights Dataset)
import seaborn as sns
df = sns.load_dataset('flights')
print(df.shape)
print(df.head(3))
(144, 3)
   year month  passengers
0  1949   Jan         112
1  1949   Feb         118
2  1949   Mar         132

What does our data look like?#

Each row is a month/year and a count of airline passengers.

  • 'year': the year
  • 'month': the month name
  • 'passengers': total airline passengers that month

For time series work, we need a pandas datetime index.

# Combine 'year' and 'month', and set as new datetime index
df['date'] = pd.to_datetime(df['year'].astype(str) + '-' + df['month'].astype(str) + '-01')
df = df.set_index('date')
print(df.head(3))
            year month  passengers
date                              
1949-01-01  1949   Jan         112
1949-02-01  1949   Feb         118
1949-03-01  1949   Mar         132
# Plotting original time series: monthly passengers
import matplotlib.pyplot as plt
df['passengers'].plot(figsize=(10, 4), title='Monthly Airline Passengers')
plt.ylabel('Passengers')
plt.xlabel('Date')
plt.show()
No description has been provided for this image

What is resampling?#

Resampling means changing the frequency of your data.

  • Downsampling: going from higher to lower frequency (monthly to yearly)
  • Upsampling: going from lower to higher frequency (monthly to daily)

You can use resampling to spot yearly or quarterly trends.

# Downsample: resample monthly data to yearly by summing passengers
annual = df['passengers'].resample('Y').sum()
print(annual.head())
date
1949-12-31    1520
1950-12-31    1676
1951-12-31    2042
1952-12-31    2364
1953-12-31    2700
Freq: YE-DEC, Name: passengers, dtype: int64
# Plot the annual passenger totals
annual.plot(kind='bar', figsize=(7,4), title='Yearly Total Passengers')
plt.ylabel('Passengers')
plt.xlabel('Year')
plt.tight_layout()
plt.show()
No description has been provided for this image
# Upsample: move from monthly to daily (with forward fill for missing data)
daily = df['passengers'].resample('D').ffill()
print(daily[:10])
date
1949-01-01    112
1949-01-02    112
1949-01-03    112
1949-01-04    112
1949-01-05    112
1949-01-06    112
1949-01-07    112
1949-01-08    112
1949-01-09    112
1949-01-10    112
Freq: D, Name: passengers, dtype: int64

Shifting data: comparing to the past or future#

Shifting means moving values forward or backward in time.

This helps us compute things like month-over-month or year-over-year changes.

# Create a new column showing previous month's passengers (shift by 1 row)
df['previous_month'] = df['passengers'].shift(1)
print(df[['passengers', 'previous_month']].head(5))
            passengers  previous_month
date                                  
1949-01-01         112             NaN
1949-02-01         118           112.0
1949-03-01         132           118.0
1949-04-01         129           132.0
1949-05-01         121           129.0
# Calculate percent change from previous month
df['percent_change'] = df['passengers'].pct_change() * 100
print(df[['passengers', 'percent_change']].head())
            passengers  percent_change
date                                  
1949-01-01         112             NaN
1949-02-01         118        5.357143
1949-03-01         132       11.864407
1949-04-01         129       -2.272727
1949-05-01         121       -6.201550
# Plot percent monthly change
df['percent_change'].plot(figsize=(10,4), grid=True, title='Percent Change vs. Previous Month')
plt.ylabel('Percent change')
plt.xlabel('Date')
plt.axhline(0, color='grey', linestyle='--')
plt.show()
No description has been provided for this image

Rolling windows: moving averages and more#

A rolling window looks at a set of recent periods (like 3 months) to summarize local trends.

Commonly, rolling windows compute averages to smooth noisy data.

# Compute rolling 12-month average of passengers
df['rolling_mean_12'] = df['passengers'].rolling(window=12).mean()
print(df[['passengers', 'rolling_mean_12']].tail(13))
            passengers  rolling_mean_12
date                                   
1959-12-01         405       428.333333
1960-01-01         417       433.083333
1960-02-01         391       437.166667
1960-03-01         419       438.250000
1960-04-01         461       443.666667
1960-05-01         472       448.000000
1960-06-01         535       453.250000
1960-07-01         622       459.416667
1960-08-01         606       463.333333
1960-09-01         508       467.083333
1960-10-01         461       471.583333
1960-11-01         390       473.916667
1960-12-01         432       476.166667
# Plot original passengers and rolling mean
plt.figure(figsize=(10,4))
plt.plot(df.index, df['passengers'], label='Monthly Passengers')
plt.plot(df.index, df['rolling_mean_12'], label='12-Month Moving Avg', linewidth=2)
plt.legend()
plt.title('Passengers vs. 12-Month Rolling Average')
plt.ylabel('Passengers')
plt.xlabel('Date')
plt.show()
No description has been provided for this image
# Rolling standard deviation (shows how much results vary month to month)
df['rolling_std_6'] = df['passengers'].rolling(window=6).std()
print(df[['passengers', 'rolling_std_6']].tail(8))
            passengers  rolling_std_6
date                                 
1960-05-01         472      32.010936
1960-06-01         535      51.685265
1960-07-01         622      83.891994
1960-08-01         606      82.470399
1960-09-01         508      67.495185
1960-10-01         461      67.495185
1960-11-01         390      87.805846
1960-12-01         432      94.200672
# Combine shifting and rolling: month-over-month change in rolling average
df['rolling_mean_diff'] = df['rolling_mean_12'].diff()
print(df[['rolling_mean_12', 'rolling_mean_diff']].tail())
            rolling_mean_12  rolling_mean_diff
date                                          
1960-08-01       463.333333           3.916667
1960-09-01       467.083333           3.750000
1960-10-01       471.583333           4.500000
1960-11-01       473.916667           2.333333
1960-12-01       476.166667           2.250000

Mini-Project: Highlighting key periods with rolling averages#

Let us combine what we have learned.

We will find the 6-month period that had the largest average number of passengers.

# Find rolling 6-month averages and highlight the period with the highest value
df['rolling_mean_6'] = df['passengers'].rolling(window=6).mean()
max_idx = df['rolling_mean_6'].idxmax()
max_period = df.loc[max_idx- pd.DateOffset(months=5):max_idx]
print(max_period[['passengers', 'rolling_mean_6']])
            passengers  rolling_mean_6
date                                  
1960-04-01         461      409.166667
1960-05-01         472      427.500000
1960-06-01         535      449.166667
1960-07-01         622      483.333333
1960-08-01         606      519.166667
1960-09-01         508      534.000000

Recap: Resampling, Shifting, and Rolling Windows in Pandas#

  • Resampling adjusts how often we measure our data.
  • Shifting lets us compare values across time.
  • Rolling windows help us smooth, spot trends, and measure volatility.

With these tools, you can handle many real-life time-series questions.

Thank you for learning with us!

Try applying rolling and shifting operations to your own datasets!

If you learned something new, please like and subscribe for more pandas tutorials on our channel.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.