Lesson 25 · Probability and Statistics in python
Understanding Regression Diagnostics: Residuals and Key Assumptions Explained
In this lesson, you will learn how to check the quality of a regression model by exploring residuals and testing model assumptions. These checks are…
- CourseProbability and Statistics in python
- Lesson25 of 35
- Video12 min
- FormatJupyter notebook · 15 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbDiagnosing Regression: Residuals and Assumptions#
In this lesson, you will learn how to check the quality of a regression model by exploring residuals and testing model assumptions.
These checks are important to know if your predictions and inferences are trustworthy.
We will explore concepts step by step, with real data, helpful visuals, and practical tips!
What is a Residual?#
A residual is the difference between the actual value and the value predicted by the regression line.
It tells you how far off your prediction was for each point.
If your model is good, residuals will show some nice characteristics, like randomness and no patterns.
import warnings
warnings.filterwarnings("ignore") # Suppress warnings for cleaner output
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns
# Data setup
mpg = sns.load_dataset("mpg")
print("MPG dataset shape:", mpg.shape)
display(mpg.head())
# Drop rows with missing values for simplicity
mpg_clean = mpg.dropna()
print("Rows after removing missing:", mpg_clean.shape[0])
Building a Simple Linear Regression#
Lets predict MPG (miles per gallon) using car weight.
First, we will look at the scatterplot and then fit a model.
# Visualize the relationship between weight and mpg
plt.figure(figsize=(7, 5))
sns.scatterplot(data=mpg_clean, x="weight", y="mpg")
plt.title("MPG vs. Weight")
plt.xlabel("Car Weight")
plt.ylabel("Miles Per Gallon (MPG)")
plt.show()
from sklearn.linear_model import LinearRegression
from sklearn.model_selection import train_test_split
# Define X and y
X = mpg_clean[["weight"]]
y = mpg_clean["mpg"]
# Train/test split for honesty
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=42)
# Fit the regression
model = LinearRegression()
model.fit(X_train, y_train)
print("Intercept:", model.intercept_)
print("Coefficient:", model.coef_[0])
# Predict and compute residuals
y_pred = model.predict(X_test)
residuals = y_test - y_pred
print("First 5 predicted values:", y_pred[:5])
print("First 5 residuals:", residuals[:5].to_numpy())
Checking Residual Assumptions#
For linear regression, we hope that residuals are:
- Normally distributed
- Have constant variance (homoscedasticity)
- Have a mean close to zero
- Show no patterns (independence)
Let us check these one by one!
# 1. Plot residuals vs predicted
plt.figure(figsize=(7,5))
plt.scatter(y_pred, residuals)
plt.axhline(0, linestyle='--', color='red')
plt.xlabel("Predicted MPG")
plt.ylabel("Residuals")
plt.title("Residuals vs. Predicted Values")
plt.show()
# 2. Check if residuals are normally distributed
sns.histplot(residuals, bins=20, kde=True)
plt.title("Residuals Histogram")
plt.xlabel("Residual")
plt.show()
# 3. Quantitative check: mean and standard deviation
print("Mean of residuals:", np.mean(residuals))
print("Standard deviation:", np.std(residuals))
# 4. Test for normality: Shapiro-Wilk test
from scipy.stats import shapiro
stat, p = shapiro(residuals)
print(f"Shapiro-Wilk test p-value: {p:.3f}")
if p > 0.05:
print("Residuals look normal.")
else:
print("Residuals do not look normal.")
# 5. Test for equal variance (homoscedasticity)
from scipy.stats import levene
split = y_pred < np.median(y_pred)
stat, p = levene(residuals[split], residuals[~split])
print(f"Levene's test for equal variance p-value: {p:.3f}")
if p > 0.05:
print("Equal variance seems reasonable.")
else:
print("May have heteroscedasticity (uneven spread).")
What Happens if Assumptions are Broken?#
If the residuals show strong patterns or are not normal, predictions or confidence intervals may not be trustworthy.
Sometimes, transformations or more advanced models can help fix these problems.
It is always a good idea to check the assumptions!
# Residuals vs input variable plot
plt.scatter(X_test["weight"], residuals)
plt.axhline(0, linestyle='--', color='red')
plt.xlabel("Weight")
plt.ylabel("Residual")
plt.title("Residuals vs. Weight")
plt.show()
# QQ plot to check for normality visually
import statsmodels.api as sm
sm.qqplot(residuals, line='45', fit=True)
plt.title("Q-Q plot for Residuals")
plt.show()
# Outlier check: Standardized residuals
from scipy.stats import zscore
z_resid = zscore(residuals)
outliers = np.where(np.abs(z_resid) > 3)[0]
print("Number of outliers (|z|>3):", len(outliers))
# Model summary with statsmodels
import statsmodels.api as sm
X_with_const = sm.add_constant(X_train)
ols_model = sm.OLS(y_train, X_with_const).fit()
print(ols_model.summary())
# Practice: Try it yourself!
col = input("Pick a numeric predictor column (e.g. 'horsepower', 'displacement', 'weight'): ")
X2 = mpg_clean[[col]]
X2_train, X2_test, y2_train, y2_test = train_test_split(X2, y, random_state=42)
model2 = LinearRegression()
model2.fit(X2_train, y2_train)
y2_pred = model2.predict(X2_test)
resid2 = y2_test - y2_pred
print("Mean residual:", np.mean(resid2))
print("Standard deviation:", np.std(resid2))
Best Practices#
- Always check your residuals
- Make visual plots, not just summary numbers
- Test your assumptions, do not just assume they are true
- Look for outliers and understand what they mean
These checks make your model stronger and more reliable!
Challenge: Try a Mini-Project!#
Can you check all assumptions with a different variable, e.g., 'horsepower' or 'acceleration'?
Try plotting, summarizing, and reporting what you find.
Share your findings in the comments section!
Recap#
- Residuals tell you about your model's mistakes.
- Good regression models have residuals that look random and normal.
- Always check model assumptions before using results.
- Use plots and numbers together for a strong analysis.
You are ready to diagnose your own regression models!
Thanks for Learning with Us!#
Do not forget to like, comment, and subscribe for more easy Python stats lessons!
See you in the next lesson!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



