Mathew K Analytics

Lesson 33 · Probability and Statistics in python

Understanding Stationarity, ACF, PACF, and ARIMA Fundamentals in Time Series Analysis

Welcome! In this lesson, you will gently learn key foundations of time series analysis in Python. We will start from the basics, explore essential concepts,…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Stationarity, ACF, PACF and ARIMA: A Beginner's Path to Time Series Analysis#

Welcome! In this lesson, you will gently learn key foundations of time series analysis in Python. We will start from the basics, explore essential concepts, and practice hands-on with real-world data.

If you have never touched ARIMA, autocorrelation, or stationarity before, do not worry - you are in the right place!

# Always start by hiding unwanted warning messages for a clear experience
import warnings
import numpy as np
np.random.seed(42)
from statsmodels.tools.sm_exceptions import ConvergenceWarning, ValueWarning
warnings.filterwarnings("ignore", category=ValueWarning)
warnings.filterwarnings("ignore", category=ConvergenceWarning)
warnings.filterwarnings("ignore", category=RuntimeWarning)
warnings.filterwarnings("ignore")

What are Time Series and Why Do They Matter?#

A time series is simply a sequence of measurements ordered by time. Examples: Daily temperatures, monthly airline passengers, stock prices.

Understanding patterns or predicting the future needs tools designed for sequences.

# Data setup: Load the classic Airline Passengers dataset
import pandas as pd
import matplotlib.pyplot as plt
import numpy as np

url = "https://raw.githubusercontent.com/jbrownlee/Datasets/master/airline-passengers.csv"
df = pd.read_csv(url, parse_dates=["Month"], index_col="Month")
print("Shape:", df.shape)
df.head()
Shape: (144, 1)
Passengers
Month
1949-01-01 112
1949-02-01 118
1949-03-01 132
1949-04-01 129
1949-05-01 121
# Plot the time series for a first visual inspection
df.plot(figsize=(10,4), legend=False, title="Monthly Airline Passengers")
plt.ylabel("Number of passengers")
plt.show()
No description has been provided for this image

What is Stationarity?#

A time series is stationary if its properties (mean, variance) do not change over time.

Why care? Stationarity is a key assumption for most forecasting methods, especially ARIMA.

Look at the chart above. Does the "airline passengers" series look stationary to you?

# Check for stationarity using a rolling mean and variance
window = 12  # 12 months = 1 year
rolling_mean = df["Passengers"].rolling(window=window).mean()
rolling_std = df["Passengers"].rolling(window=window).std()

plt.figure(figsize=(10,4))
plt.plot(df["Passengers"], label="Original")
plt.plot(rolling_mean, label="Rolling Mean (1 year)")
plt.plot(rolling_std, label="Rolling Std Dev (1 year)")
plt.legend()
plt.title("Stationarity: Rolling Statistics")
plt.show()
No description has been provided for this image
# The Augmented Dickey-Fuller Test: a formal test for stationarity
from statsmodels.tsa.stattools import adfuller
result = adfuller(df["Passengers"])
print("ADF Statistic:", result[0])
print("p-value:", result[1])
ADF Statistic: 0.8153688792060482
p-value: 0.991880243437641

How to Make a Series Stationary: Differencing#

Many series need to be transformed to become stationary.

Differencing means subtracting each value by its previous value. This removes trends and seasonality in many cases.

Let's try this and see what changes.

# Apply first-order differencing to remove trend
diff = df["Passengers"].diff().dropna()
plt.figure(figsize=(10,4))
plt.plot(diff, label="First-differenced series")
plt.legend()
plt.title("After First Differencing")
plt.show()
No description has been provided for this image
# Re-run the ADF test on the differenced series
result_diff = adfuller(diff)
print("ADF Statistic after differencing:", result_diff[0])
print("p-value after differencing:", result_diff[1])
ADF Statistic after differencing: -2.8292668241699994
p-value after differencing: 0.0542132902838255

What are ACF and PACF?#

ACF: Autocorrelation Function shows how today's value is related to past values. PACF: Partial Autocorrelation Function removes indirect effects, showing only direct relationships.

We use these to choose AR, MA, and ARIMA models.

# Plot ACF and PACF for the differenced series
from statsmodels.graphics.tsaplots import plot_acf, plot_pacf
fig, axes = plt.subplots(1,2, figsize=(12,4))
plot_acf(diff, ax=axes[0], lags=24, title="ACF - Airline Passengers")
plot_pacf(diff, ax=axes[1], lags=24, title="PACF - Airline Passengers")
plt.tight_layout()
plt.show()
No description has been provided for this image

Basics of ARIMA Models#

ARIMA stands for AutoRegressive Integrated Moving Average.

It combines:

  • AR: Autoregression (uses past values)
  • I: Integrated (differencing, as above)
  • MA: Moving Average (error smoothing)

The three numbers (p, d, q) in ARIMA indicate:

  • p: Number of autoregressive terms (check PACF)
  • d: Difference order (how many times you difference)
  • q: Moving average terms (check ACF)

Let's try a first ARIMA model soon!

# Fit ARIMA (p=1, d=1, q=1) to the data and visualize prediction
from statsmodels.tsa.arima.model import ARIMA
import warnings
from statsmodels.tools.sm_exceptions import ConvergenceWarning, ValueWarning
warnings.filterwarnings("ignore", category=ValueWarning)
warnings.filterwarnings("ignore", category=ConvergenceWarning)
warnings.filterwarnings("ignore", category=RuntimeWarning)

# Ensure the index has proper frequency for ARIMA
df_arima = df.copy()
df_arima.index.freq = "MS"

try:
    model = ARIMA(df_arima["Passengers"], order=(1, 1, 1))
    model_fit = model.fit()
    print(model_fit.summary())

    forecast = model_fit.predict(start=len(df_arima)-12, end=len(df_arima)-1, typ="levels")
    df_arima["Passengers"].iloc[-24:].plot(label="Actual", marker="o")
    forecast.plot(label="Forecast", marker="x")
    plt.title("ARIMA Forecast vs Last 24 Months")
    plt.legend()
    plt.show()
except Exception as e:
    print("An error occurred during ARIMA fitting or forecasting:", e)
    
                               SARIMAX Results                                
==============================================================================
Dep. Variable:             Passengers   No. Observations:                  144
Model:                 ARIMA(1, 1, 1)   Log Likelihood                -694.341
Date:                Thu, 11 Sep 2025   AIC                           1394.683
Time:                        14:40:16   BIC                           1403.571
Sample:                    01-01-1949   HQIC                          1398.294
                         - 12-01-1960                                         
Covariance Type:                  opg                                         
==============================================================================
                 coef    std err          z      P>|z|      [0.025      0.975]
------------------------------------------------------------------------------
ar.L1         -0.4742      0.123     -3.847      0.000      -0.716      -0.233
ma.L1          0.8635      0.078     11.051      0.000       0.710       1.017
sigma2       961.9270    107.433      8.954      0.000     751.362    1172.492
===================================================================================
Ljung-Box (L1) (Q):                   0.21   Jarque-Bera (JB):                 2.14
Prob(Q):                              0.65   Prob(JB):                         0.34
Heteroskedasticity (H):               7.00   Skew:                            -0.21
Prob(H) (two-sided):                  0.00   Kurtosis:                         3.43
===================================================================================

Warnings:
[1] Covariance matrix calculated using the outer product of gradients (complex-step).
No description has been provided for this image
# Handling missing data in time series
df_missing = df.copy()
df_missing.iloc[10:15] = None  # Simulate missing values
print(df_missing.isnull().sum())

# Fill using forward-fill (propagate last value)
df_missing_ffill = df_missing.fillna(method="ffill")
df_missing_ffill.head(20)
Passengers    5
dtype: int64
Passengers
Month
1949-01-01 112.0
1949-02-01 118.0
1949-03-01 132.0
1949-04-01 129.0
1949-05-01 121.0
1949-06-01 135.0
1949-07-01 148.0
1949-08-01 148.0
1949-09-01 136.0
1949-10-01 119.0
1949-11-01 119.0
1949-12-01 119.0
1950-01-01 119.0
1950-02-01 119.0
1950-03-01 119.0
1950-04-01 135.0
1950-05-01 125.0
1950-06-01 149.0
1950-07-01 170.0
1950-08-01 170.0
# Practice: Identify if a new series needs differencing
import numpy as np
np.random.seed(42)
random_walk = np.cumsum(np.random.randn(100))
plt.plot(random_walk)
plt.title("Random Walk Example")
plt.show()
No description has been provided for this image
# Check ADF test for stationarity of the random walk
from statsmodels.tsa.stattools import adfuller
rw_result = adfuller(random_walk)
print("Random walk p-value:", rw_result[1])
Random walk p-value: 0.6020814791099101
# Quick You: Choose which transformation makes data stationary
choice = input("If a series drifts upward over time, what is your first stationarizing step? (a) take log, (b) difference, (c) z-score: ")
if choice.lower().startswith("b"):
    print("Correct! Differencing is the main way to remove trends.")
else:
    print("Nice try! Differencing is usually first. Logs help with spread, z-scores scale the series.")
    
Correct! Differencing is the main way to remove trends.
# Challenge: Try ARIMA on your own synthetic data
synthetic = np.sin(np.linspace(0, 20 * np.pi, 120)) + np.random.normal(scale=0.5, size=120)
plt.plot(synthetic)
plt.title("Synthetic Time Series")
plt.show()

from statsmodels.tsa.arima.model import ARIMA
model2 = ARIMA(synthetic, order=(1,0,0))
fit2 = model2.fit()
pred = fit2.predict(start=100, end=119)
plt.plot(synthetic, label="Actual")
plt.plot(range(100,120), pred, label="ARIMA Forecast", linestyle='--')
plt.legend()
plt.title("Forecast: ARIMA on Synthetic Series")
plt.show()
No description has been provided for this image
No description has been provided for this image

Extra Tips:#

  1. Always visualize your data first.
  2. Check for missing values before modeling.
  3. Try multiple differencing levels if stationarity is uncertain.
  4. Use ACF/PACF plots as model-guides, not strict rules.
  5. Save your notebook often!

Ready for more? Try changing ARIMA settings or exploring the Sunspots dataset next.

# Mini-project (Part 1): Simulate a real-world process
np.random.seed(0)
visits_base = 100 + np.arange(60) * 2  # upward trend
visits_season = 10 * np.sin(np.linspace(0, 6*np.pi, 60))
noise = np.random.normal(0, 3, size=60)
website_visits = visits_base + visits_season + noise
plt.plot(website_visits)
plt.title("Simulated Monthly Website Visits")
plt.show()
No description has been provided for this image
# Mini-project (Part 2): Analyze stationarity for simulated visits
plt.figure(figsize=(10,4))
plt.plot(website_visits, label="Original")
plt.plot(pd.Series(website_visits).rolling(12).mean(), label="Mean")
plt.plot(pd.Series(website_visits).rolling(12).std(), label="Std Dev")
plt.legend()
plt.title("Rolling Stats of Simulated Visits")
plt.show()

result = adfuller(website_visits)
print("ADF Statistic:", result[0], "p-value:", result[1])
No description has been provided for this image
ADF Statistic: -0.02760199283161892 p-value: 0.9561890191760988

Recap: Key Lessons and Where to Go Next#

You learned:

  • How to plot and inspect a time series
  • What stationarity means and why it matters
  • How to use rolling averages and the ADF test
  • When and why to difference your data
  • Basics of ACF and PACF plots for ARIMA
  • How to fit ARIMA and check forecasts
  • How to handle missing time series values
  • Where these skills apply (from web visits to weather)

Explore, experiment, and most importantly: have fun learning time series!

Thanks for Learning! Share Your Results and Subscribe#

Thanks for studying Stationarity and ARIMA basics with us.

Try the challenge exercises on your own data.

If you liked this lesson, subscribe to our channel and share your graphs! See you next time for more hands-on data science.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.