Lesson 26 · Python For Time Series
Understanding Model Diagnostics and Residual Analysis for Time Series Forecasting in Python
In this lesson, we will learn how to check how well a model fits data, spot problems, and improve predictions using residual analysis. You will use…
- CoursePython For Time Series
- Lesson26 of 30
- Video14 min
- FormatJupyter notebook · 20 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynb
Model Diagnostics and Residual Analysis in Python#
In this lesson, we will learn how to check how well a model fits data, spot problems, and improve predictions using residual analysis.
You will use real-world time series data, explore important plots, and understand key ideas step by step.
Let us get started!
# Basic setup: Import core libraries
import warnings; warnings.filterwarnings("ignore")
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
# Data setup: Load a classic time series (Airline Passengers)
df = pd.read_csv(
"https://raw.githubusercontent.com/jbrownlee/Datasets/master/airline-passengers.csv",
parse_dates=['Month']
)
df = df.rename(columns={'Passengers': 'y', 'Month': 'date'})
df['date'] = pd.to_datetime(df['date'])
df = df.set_index('date')
print('Data shape:', df.shape)
df.head()
# Plot the time series
df['y'].plot(figsize=(8, 4), title='Airline Passengers Over Time')
plt.xlabel('Date')
plt.ylabel('Number of Passengers')
plt.show()
What are Residuals?#
Residuals are the leftover parts after a model makes predictions.
They help show what the model missed.
If they look random, it is good. If there are patterns, we can do better.
# Build a basic forecasting model: moving average
df['y_pred'] = df['y'].rolling(window=12).mean().shift(1)
plt.figure(figsize=(8,4))
plt.plot(df.index, df['y'], label='Actual')
plt.plot(df.index, df['y_pred'], label='Prediction', linestyle='--')
plt.legend()
plt.title('Moving Average Prediction vs Actual')
plt.show()
# Calculate residuals
df['residual'] = df['y'] - df['y_pred']
plt.figure(figsize=(8,3))
plt.plot(df.index, df['residual'], color='purple')
plt.axhline(0, color='black', linestyle='--')
plt.title('Residuals Over Time')
plt.ylabel('Residual (Error)')
plt.show()
Why Plot Residuals?#
Plotting residuals can reveal bias, missed patterns, or data issues.
If you see shapes, trends, or repeating errors, you may want to try a better model.
# Remove missing predictions (first 12 months)
df_valid = df.dropna(subset=['y_pred'])
# Check a simple summary
print('Mean residual (should be close to zero):', round(df_valid['residual'].mean(), 2))
print('Standard deviation of residuals:', round(df_valid['residual'].std(), 2))
# Histogram: How do errors behave?
plt.hist(df_valid['residual'], bins=20, color='orange', edgecolor='black')
plt.title('Distribution of Residuals')
plt.xlabel('Residual Value')
plt.ylabel('Count')
plt.show()
# Residuals vs predictions: Any patterns?
plt.scatter(df_valid['y_pred'], df_valid['residual'], color='green', alpha=0.6)
plt.axhline(0, color='black', linestyle='--')
plt.xlabel('Predicted Value')
plt.ylabel('Residual (Error)')
plt.title('Residuals vs Predictions')
plt.show()
Model Diagnostics: Key Questions#
- Are the errors random or do they show patterns?
- Are errors bigger at certain times or sizes?
- Did we miss seasonality or trends?
- Can another model do better?
Let us test for randomness next.
# Check autocorrelation: Are errors independent?
from pandas.plotting import autocorrelation_plot
autocorrelation_plot(df_valid['residual'])
plt.title('Residual Autocorrelation')
plt.show()
# Q-Q plot: Do errors look normal?
import scipy.stats as stats
stats.probplot(df_valid['residual'], dist='norm', plot=plt)
plt.title('Q-Q Plot of Residuals')
plt.show()
# Try another model: simple ARIMA
from statsmodels.tsa.arima.model import ARIMA
model = ARIMA(df['y'], order=(1,1,1))
model_fit = model.fit()
df['arima_pred'] = model_fit.predict(start=1, end=len(df))
df['arima_resid'] = df['y'] - df['arima_pred']
plt.figure(figsize=(8,3))
plt.plot(df.index, df['arima_resid'], color='red', label='ARIMA Residuals')
plt.axhline(0, color='black', linestyle='--')
plt.title('ARIMA Model Residuals')
plt.legend()
plt.show()
# Compare: Histogram of both models' errors
plt.hist(df_valid['residual'], bins=16, alpha=0.5, label='Moving Avg')
plt.hist(df['arima_resid'].dropna(), bins=16, alpha=0.5, label='ARIMA')
plt.xlabel('Residual Value')
plt.ylabel('Count')
plt.title('Comparison: Residuals Distribution')
plt.legend()
plt.show()
# Input: Try your own number for passenger count!
your_number = int(input("Guess a passenger count from the chart above and type it here: "))
print("You entered:", your_number)
if your_number > df['y'].max():
print("That number is higher than any real month!")
elif your_number < df['y'].min():
print("That is a small number for this dataset!")
else:
print("That is a realistic guess.")
Troubleshooting Tips#
- If errors are not random, change models or add features.
- Always look at charts before trusting results.
- If predictions look behind or ahead, check your code for time shifts.
- Outliers can trick your model. Mark or remove them if needed.
- Missing values give bad errors. Use dropna() if you see NaNs.
# Mini-project Part 1: Residual analysis for another dataset
df2 = pd.read_csv(
"https://raw.githubusercontent.com/jbrownlee/Datasets/master/shampoo.csv",
parse_dates=['Month']
)
df2["Month"] = pd.date_range(start="1901-01", periods=len(df2), freq="M")
df2["Month"] = pd.to_datetime(df2["Month"], format="%y-%m")
df2 = df2.rename(columns={'Sales': 'y', 'Month': 'date'})
df2['date'] = pd.to_datetime(df2['date'])
df2 = df2.set_index('date')
df2['y_pred'] = df2['y'].rolling(window=6).mean().shift(1)
df2['residual'] = df2['y'] - df2['y_pred']
plt.plot(df2.index, df2['residual'], color='teal')
plt.axhline(0, color='black', linestyle='--')
plt.title('Shampoo Sales: Residuals (6-month mean)')
plt.tight_layout()
plt.show()
# Mini-project Part 2: Residual summary and improvement hint
print('Shampoo model residual mean:', round(df2['residual'].mean(),2))
print('Std of residuals:', round(df2['residual'].std(),2))
if abs(df2['residual'].mean()) > 15:
print('Try a different window size for moving average or another model.')
else:
print('Nice! Residuals are centered. The model catches the pattern fairly well.')
# Useful shortcut: Residual plotting function
def plot_resids(y, y_pred, title):
res = y - y_pred
plt.plot(res, color='navy')
plt.axhline(0, color='gray', linestyle='--')
plt.title(title)
plt.show()
# Example use
plot_resids(df2['y'], df2['y_pred'], 'Shampoo: Quick Residuals')
# Challenge: Try your own moving average window
window_size = int(input("Pick a window size (example: 3, 6, or 12): "))
df2['y_pred_custom'] = df2['y'].rolling(window=window_size).mean().shift(1)
plot_resids(df2['y'], df2['y_pred_custom'], f'Shampoo: Residuals (window={window_size})')
print('Did you see smaller or bigger residuals this time?')
Recap#
- You learned how to check model fit using residuals.
- You saw that looking for random errors means your model is on track.
- You used simple models and a classic ARIMA on real data.
- You learned to spot problems and fix them with new ideas.
Keep practicing and let us know your questions!
Thanks for learning Model Diagnostics!#
If you found this lesson helpful, please like, comment, and subscribe for more Python tutorials.
Happy coding!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



