Mathew K Analytics

Lesson 26 · Python For Time Series

Understanding Model Diagnostics and Residual Analysis for Time Series Forecasting in Python

In this lesson, we will learn how to check how well a model fits data, spot problems, and improve predictions using residual analysis. You will use…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb
 

Model Diagnostics and Residual Analysis in Python#

In this lesson, we will learn how to check how well a model fits data, spot problems, and improve predictions using residual analysis.

You will use real-world time series data, explore important plots, and understand key ideas step by step.

Let us get started!

# Basic setup: Import core libraries
import warnings; warnings.filterwarnings("ignore")
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
 
 
# Data setup: Load a classic time series (Airline Passengers)
df = pd.read_csv(
    "https://raw.githubusercontent.com/jbrownlee/Datasets/master/airline-passengers.csv",
    parse_dates=['Month']
)
df = df.rename(columns={'Passengers': 'y', 'Month': 'date'})
df['date'] = pd.to_datetime(df['date'])
df = df.set_index('date')
print('Data shape:', df.shape)
df.head()
Data shape: (144, 1)
y
date
1949-01-01 112
1949-02-01 118
1949-03-01 132
1949-04-01 129
1949-05-01 121
# Plot the time series
df['y'].plot(figsize=(8, 4), title='Airline Passengers Over Time')
plt.xlabel('Date')
plt.ylabel('Number of Passengers')
plt.show()
 
 
No description has been provided for this image

What are Residuals?#

Residuals are the leftover parts after a model makes predictions.

They help show what the model missed.

If they look random, it is good. If there are patterns, we can do better.

# Build a basic forecasting model: moving average
df['y_pred'] = df['y'].rolling(window=12).mean().shift(1)
 
plt.figure(figsize=(8,4))
plt.plot(df.index, df['y'], label='Actual')
plt.plot(df.index, df['y_pred'], label='Prediction', linestyle='--')
plt.legend()
plt.title('Moving Average Prediction vs Actual')
plt.show()
No description has been provided for this image
# Calculate residuals
df['residual'] = df['y'] - df['y_pred']

plt.figure(figsize=(8,3))
plt.plot(df.index, df['residual'], color='purple')
plt.axhline(0, color='black', linestyle='--')
plt.title('Residuals Over Time')
plt.ylabel('Residual (Error)')
plt.show()
No description has been provided for this image

Why Plot Residuals?#

Plotting residuals can reveal bias, missed patterns, or data issues.

If you see shapes, trends, or repeating errors, you may want to try a better model.

# Remove missing predictions (first 12 months)
df_valid = df.dropna(subset=['y_pred'])

# Check a simple summary
print('Mean residual (should be close to zero):', round(df_valid['residual'].mean(), 2))
print('Standard deviation of residuals:', round(df_valid['residual'].std(), 2))
Mean residual (should be close to zero): 17.59
Standard deviation of residuals: 46.69
# Histogram: How do errors behave?
plt.hist(df_valid['residual'], bins=20, color='orange', edgecolor='black')
plt.title('Distribution of Residuals')
plt.xlabel('Residual Value')
plt.ylabel('Count')
plt.show()
No description has been provided for this image
# Residuals vs predictions: Any patterns?
plt.scatter(df_valid['y_pred'], df_valid['residual'], color='green', alpha=0.6)
plt.axhline(0, color='black', linestyle='--')
plt.xlabel('Predicted Value')
plt.ylabel('Residual (Error)')
plt.title('Residuals vs Predictions')
plt.show()
No description has been provided for this image

Model Diagnostics: Key Questions#

  1. Are the errors random or do they show patterns?
  2. Are errors bigger at certain times or sizes?
  3. Did we miss seasonality or trends?
  4. Can another model do better?

Let us test for randomness next.

# Check autocorrelation: Are errors independent?
from pandas.plotting import autocorrelation_plot
autocorrelation_plot(df_valid['residual'])
plt.title('Residual Autocorrelation')
plt.show()
No description has been provided for this image
# Q-Q plot: Do errors look normal?
import scipy.stats as stats
stats.probplot(df_valid['residual'], dist='norm', plot=plt)
plt.title('Q-Q Plot of Residuals')
plt.show()
No description has been provided for this image
# Try another model: simple ARIMA
from statsmodels.tsa.arima.model import ARIMA

model = ARIMA(df['y'], order=(1,1,1))
model_fit = model.fit()
df['arima_pred'] = model_fit.predict(start=1, end=len(df))
df['arima_resid'] = df['y'] - df['arima_pred']

plt.figure(figsize=(8,3))
plt.plot(df.index, df['arima_resid'], color='red', label='ARIMA Residuals')
plt.axhline(0, color='black', linestyle='--')
plt.title('ARIMA Model Residuals')
plt.legend()
plt.show()
No description has been provided for this image
# Compare: Histogram of both models' errors
plt.hist(df_valid['residual'], bins=16, alpha=0.5, label='Moving Avg')
plt.hist(df['arima_resid'].dropna(), bins=16, alpha=0.5, label='ARIMA')
plt.xlabel('Residual Value')
plt.ylabel('Count')
plt.title('Comparison: Residuals Distribution')
plt.legend()
plt.show()
No description has been provided for this image
# Input: Try your own number for passenger count!
your_number = int(input("Guess a passenger count from the chart above and type it here: "))
print("You entered:", your_number)
if your_number > df['y'].max():
    print("That number is higher than any real month!")
elif your_number < df['y'].min():
    print("That is a small number for this dataset!")
else:
    print("That is a realistic guess.")
    
You entered: 400
That is a realistic guess.
 

Troubleshooting Tips#

  1. If errors are not random, change models or add features.
  2. Always look at charts before trusting results.
  3. If predictions look behind or ahead, check your code for time shifts.
  4. Outliers can trick your model. Mark or remove them if needed.
  5. Missing values give bad errors. Use dropna() if you see NaNs.
# Mini-project Part 1: Residual analysis for another dataset
df2 = pd.read_csv(
    "https://raw.githubusercontent.com/jbrownlee/Datasets/master/shampoo.csv",
    parse_dates=['Month']
)
df2["Month"] = pd.date_range(start="1901-01", periods=len(df2), freq="M")
df2["Month"] = pd.to_datetime(df2["Month"], format="%y-%m")
df2 = df2.rename(columns={'Sales': 'y', 'Month': 'date'})
df2['date'] = pd.to_datetime(df2['date'])
df2 = df2.set_index('date')
df2['y_pred'] = df2['y'].rolling(window=6).mean().shift(1)
df2['residual'] = df2['y'] - df2['y_pred']

plt.plot(df2.index, df2['residual'], color='teal')
plt.axhline(0, color='black', linestyle='--')
plt.title('Shampoo Sales: Residuals (6-month mean)')
plt.tight_layout()
plt.show()
No description has been provided for this image
# Mini-project Part 2: Residual summary and improvement hint
print('Shampoo model residual mean:', round(df2['residual'].mean(),2))
print('Std of residuals:', round(df2['residual'].std(),2))
if abs(df2['residual'].mean()) > 15:
    print('Try a different window size for moving average or another model.')
else:
    print('Nice! Residuals are centered. The model catches the pattern fairly well.')
    
Shampoo model residual mean: 47.94
Std of residuals: 75.38
Try a different window size for moving average or another model.
# Useful shortcut: Residual plotting function
def plot_resids(y, y_pred, title):
    res = y - y_pred
    plt.plot(res, color='navy')
    plt.axhline(0, color='gray', linestyle='--')
    plt.title(title)
    plt.show()

# Example use
plot_resids(df2['y'], df2['y_pred'], 'Shampoo: Quick Residuals')
No description has been provided for this image
# Challenge: Try your own moving average window
window_size = int(input("Pick a window size (example: 3, 6, or 12): "))
df2['y_pred_custom'] = df2['y'].rolling(window=window_size).mean().shift(1)
plot_resids(df2['y'], df2['y_pred_custom'], f'Shampoo: Residuals (window={window_size})')

print('Did you see smaller or bigger residuals this time?')
No description has been provided for this image
Did you see smaller or bigger residuals this time?
 

Recap#

  • You learned how to check model fit using residuals.
  • You saw that looking for random errors means your model is on track.
  • You used simple models and a classic ARIMA on real data.
  • You learned to spot problems and fix them with new ideas.

Keep practicing and let us know your questions!

Thanks for learning Model Diagnostics!#

If you found this lesson helpful, please like, comment, and subscribe for more Python tutorials.

Happy coding!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.