Mathew K Analytics

Lesson 25 · Probability and Statistics in python

Understanding Regression Diagnostics: Residuals and Key Assumptions Explained

In this lesson, you will learn how to check the quality of a regression model by exploring residuals and testing model assumptions. These checks are…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Diagnosing Regression: Residuals and Assumptions#

In this lesson, you will learn how to check the quality of a regression model by exploring residuals and testing model assumptions.

These checks are important to know if your predictions and inferences are trustworthy.

We will explore concepts step by step, with real data, helpful visuals, and practical tips!

What is a Residual?#

A residual is the difference between the actual value and the value predicted by the regression line.

It tells you how far off your prediction was for each point.

If your model is good, residuals will show some nice characteristics, like randomness and no patterns.

import warnings
warnings.filterwarnings("ignore")  # Suppress warnings for cleaner output

import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns

# Data setup
mpg = sns.load_dataset("mpg")
print("MPG dataset shape:", mpg.shape)
display(mpg.head())
MPG dataset shape: (398, 9)
mpg cylinders displacement horsepower weight acceleration model_year origin name
0 18.0 8 307.0 130.0 3504 12.0 70 usa chevrolet chevelle malibu
1 15.0 8 350.0 165.0 3693 11.5 70 usa buick skylark 320
2 18.0 8 318.0 150.0 3436 11.0 70 usa plymouth satellite
3 16.0 8 304.0 150.0 3433 12.0 70 usa amc rebel sst
4 17.0 8 302.0 140.0 3449 10.5 70 usa ford torino
# Drop rows with missing values for simplicity
mpg_clean = mpg.dropna()
print("Rows after removing missing:", mpg_clean.shape[0])
Rows after removing missing: 392

Building a Simple Linear Regression#

Lets predict MPG (miles per gallon) using car weight.

First, we will look at the scatterplot and then fit a model.

# Visualize the relationship between weight and mpg
plt.figure(figsize=(7, 5))
sns.scatterplot(data=mpg_clean, x="weight", y="mpg")
plt.title("MPG vs. Weight")
plt.xlabel("Car Weight")
plt.ylabel("Miles Per Gallon (MPG)")
plt.show()
No description has been provided for this image
from sklearn.linear_model import LinearRegression
from sklearn.model_selection import train_test_split

# Define X and y
X = mpg_clean[["weight"]]
y = mpg_clean["mpg"]

# Train/test split for honesty
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=42)

# Fit the regression
model = LinearRegression()
model.fit(X_train, y_train)

print("Intercept:", model.intercept_)
print("Coefficient:", model.coef_[0])
Intercept: 47.283253138620495
Coefficient: -0.007938232270314252
# Predict and compute residuals
y_pred = model.predict(X_test)
residuals = y_test - y_pred

print("First 5 predicted values:", y_pred[:5])
print("First 5 residuals:", residuals[:5].to_numpy())
First 5 predicted values: [29.9064627  25.09589394 32.99443505 31.76400905 25.1355851 ]
First 5 residuals: [-3.9064627  -3.49589394  3.10556495 -5.76400905  1.8644149 ]

Checking Residual Assumptions#

For linear regression, we hope that residuals are:

  • Normally distributed
  • Have constant variance (homoscedasticity)
  • Have a mean close to zero
  • Show no patterns (independence)

Let us check these one by one!

# 1. Plot residuals vs predicted
plt.figure(figsize=(7,5))
plt.scatter(y_pred, residuals)
plt.axhline(0, linestyle='--', color='red')
plt.xlabel("Predicted MPG")
plt.ylabel("Residuals")
plt.title("Residuals vs. Predicted Values")
plt.show()
No description has been provided for this image
# 2. Check if residuals are normally distributed
sns.histplot(residuals, bins=20, kde=True)
plt.title("Residuals Histogram")
plt.xlabel("Residual")
plt.show()
No description has been provided for this image
# 3. Quantitative check: mean and standard deviation
print("Mean of residuals:", np.mean(residuals))
print("Standard deviation:", np.std(residuals))
Mean of residuals: -0.802319667438032
Standard deviation: 4.009479107636392
# 4. Test for normality: Shapiro-Wilk test
from scipy.stats import shapiro
stat, p = shapiro(residuals)
print(f"Shapiro-Wilk test p-value: {p:.3f}")
if p > 0.05:
    print("Residuals look normal.")
else:
    print("Residuals do not look normal.")
    
Shapiro-Wilk test p-value: 0.018
Residuals do not look normal.
# 5. Test for equal variance (homoscedasticity)
from scipy.stats import levene
split = y_pred < np.median(y_pred)
stat, p = levene(residuals[split], residuals[~split])
print(f"Levene's test for equal variance p-value: {p:.3f}")
if p > 0.05:
    print("Equal variance seems reasonable.")
else:
    print("May have heteroscedasticity (uneven spread).")
    
Levene's test for equal variance p-value: 0.017
May have heteroscedasticity (uneven spread).

What Happens if Assumptions are Broken?#

If the residuals show strong patterns or are not normal, predictions or confidence intervals may not be trustworthy.

Sometimes, transformations or more advanced models can help fix these problems.

It is always a good idea to check the assumptions!

# Residuals vs input variable plot
plt.scatter(X_test["weight"], residuals)
plt.axhline(0, linestyle='--', color='red')
plt.xlabel("Weight")
plt.ylabel("Residual")
plt.title("Residuals vs. Weight")
plt.show()
No description has been provided for this image
# QQ plot to check for normality visually
import statsmodels.api as sm
sm.qqplot(residuals, line='45', fit=True)
plt.title("Q-Q plot for Residuals")
plt.show()
No description has been provided for this image
# Outlier check: Standardized residuals
from scipy.stats import zscore
z_resid = zscore(residuals)
outliers = np.where(np.abs(z_resid) > 3)[0]
print("Number of outliers (|z|>3):", len(outliers))
Number of outliers (|z|>3): 1
# Model summary with statsmodels
import statsmodels.api as sm
X_with_const = sm.add_constant(X_train)
ols_model = sm.OLS(y_train, X_with_const).fit()
print(ols_model.summary())
                            OLS Regression Results                            
==============================================================================
Dep. Variable:                    mpg   R-squared:                       0.696
Model:                            OLS   Adj. R-squared:                  0.695
Method:                 Least Squares   F-statistic:                     667.6
Date:                Wed, 03 Sep 2025   Prob (F-statistic):           2.03e-77
Time:                        09:14:51   Log-Likelihood:                -853.55
No. Observations:                 294   AIC:                             1711.
Df Residuals:                     292   BIC:                             1718.
Df Model:                           1                                         
Covariance Type:            nonrobust                                         
==============================================================================
                 coef    std err          t      P>|t|      [0.025      0.975]
------------------------------------------------------------------------------
const         47.2833      0.949     49.832      0.000      45.416      49.151
weight        -0.0079      0.000    -25.838      0.000      -0.009      -0.007
==============================================================================
Omnibus:                       26.720   Durbin-Watson:                   2.198
Prob(Omnibus):                  0.000   Jarque-Bera (JB):               35.048
Skew:                           0.654   Prob(JB):                     2.45e-08
Kurtosis:                       4.072   Cond. No.                     1.14e+04
==============================================================================

Notes:
[1] Standard Errors assume that the covariance matrix of the errors is correctly specified.
[2] The condition number is large, 1.14e+04. This might indicate that there are
strong multicollinearity or other numerical problems.
# Practice: Try it yourself!
col = input("Pick a numeric predictor column (e.g. 'horsepower', 'displacement', 'weight'): ")
X2 = mpg_clean[[col]]
X2_train, X2_test, y2_train, y2_test = train_test_split(X2, y, random_state=42)
model2 = LinearRegression()
model2.fit(X2_train, y2_train)
y2_pred = model2.predict(X2_test)
resid2 = y2_test - y2_pred
print("Mean residual:", np.mean(resid2))
print("Standard deviation:", np.std(resid2))
Mean residual: -0.9872450811840342
Standard deviation: 4.575426580543753

Best Practices#

  • Always check your residuals
  • Make visual plots, not just summary numbers
  • Test your assumptions, do not just assume they are true
  • Look for outliers and understand what they mean

These checks make your model stronger and more reliable!

Challenge: Try a Mini-Project!#

Can you check all assumptions with a different variable, e.g., 'horsepower' or 'acceleration'?

Try plotting, summarizing, and reporting what you find.

Share your findings in the comments section!

Recap#

  • Residuals tell you about your model's mistakes.
  • Good regression models have residuals that look random and normal.
  • Always check model assumptions before using results.
  • Use plots and numbers together for a strong analysis.

You are ready to diagnose your own regression models!

Thanks for Learning with Us!#

Do not forget to like, comment, and subscribe for more easy Python stats lessons!

See you in the next lesson!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.