Mathew K Analytics

Lesson 11 · Statistics for data analysts

Simple Linear Regression Explained | Statistics #11

Video eleven of the 15-part series: moving from comparing groups to modeling a real relationship directly. Real mtcars data: predicting fuel economy from…

What you'll learn

Datasets used in this lesson

Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.

📓 Full notebook

Download .ipynb

Statistics for Data Analysts, Video 11: Simple Linear Regression#

  • Video eleven of the 15-part series: moving from comparing groups to modeling a real relationship directly.
  • Real mtcars data: predicting fuel economy from weight.
  • Let's get into it.

Before You Start#

  • Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
  • You'll need pandas, NumPy, SciPy, and Matplotlib.
  • Place mtcars.csv in the same folder as this notebook.
import pandas as pd
import numpy as np
from scipy import stats
import matplotlib.pyplot as plt

mtcars = pd.read_csv('mtcars.csv')
print(mtcars[['model', 'wt', 'mpg']].head())
               model     wt   mpg
0          Mazda RX4  2.620  21.0
1      Mazda RX4 Wag  2.875  21.0
2         Datsun 710  2.320  22.8
3     Hornet 4 Drive  3.215  21.4
4  Hornet Sportabout  3.440  18.7

Part 1: Fitting the Model#

x = mtcars['wt']
y = mtcars['mpg']
result = stats.linregress(x, y)
print(f'Real slope: {result.slope:.3f}')
print(f'Real intercept: {result.intercept:.3f}')
print(f'Real p-value: {result.pvalue:.2e}')
Real slope: -5.344
Real intercept: 37.285
Real p-value: 1.29e-10

The real slope, about -5.34, means: for every additional 1,000 pounds of real weight, this model predicts about 5.34 fewer real miles per gallon. The real intercept, about 37.3, is the model's predicted mpg at zero weight, mathematically necessary for the line but not a meaningful real car. The real p-value is far below 0.05: strong evidence the real relationship isn't just noise.

plt.figure(figsize=(8, 5))
plt.scatter(x, y, color='steelblue', label='Real cars')
x_line = np.linspace(x.min(), x.max(), 100)
plt.plot(x_line, result.intercept + result.slope * x_line, color='darkred', linewidth=2, label='Fitted line')
plt.title('Real mtcars: Weight vs. MPG')
plt.xlabel('Weight (1000 lbs)')
plt.ylabel('Miles per Gallon')
plt.legend()
plt.show()
No description has been provided for this image

Part 2: R-Squared, and What It Actually Measures#

r_squared = result.rvalue ** 2
print(f'Real correlation (r): {result.rvalue:.4f}')
print(f'Real R-squared: {r_squared:.4f}')
print(f'Interpretation: weight explains about {r_squared:.1%} of the real variance in mpg')
Real correlation (r): -0.8677
Real R-squared: 0.7528
Interpretation: weight explains about 75.3% of the real variance in mpg

Part 3: Predictions and Prediction Intervals#

n = len(x)
yhat = result.intercept + result.slope * x
residuals = y - yhat
mse = np.sum(residuals ** 2) / (n - 2)
se_resid = np.sqrt(mse)
x_mean = x.mean()
sxx = np.sum((x - x_mean) ** 2)
print(f'Real residual standard error: {se_resid:.3f}')
Real residual standard error: 3.046
x0 = 3.0
pred_mpg = result.intercept + result.slope * x0
se_pred = se_resid * np.sqrt(1 + 1 / n + (x0 - x_mean) ** 2 / sxx)
t_crit = stats.t.ppf(0.975, df=n - 2)
pi_low, pi_high = pred_mpg - t_crit * se_pred, pred_mpg + t_crit * se_pred
print(f'For a real 3,000 lb car, predicted mpg: {pred_mpg:.2f}')
print(f'Real 95% prediction interval: ({pi_low:.2f}, {pi_high:.2f})')
For a real 3,000 lb car, predicted mpg: 21.25
Real 95% prediction interval: (14.93, 27.57)

Part 4: Residual Diagnostics#

fig, axes = plt.subplots(1, 2, figsize=(11, 4))
axes[0].scatter(yhat, residuals, color='steelblue')
axes[0].axhline(0, color='darkred', linestyle='--')
axes[0].set_title('Real Residuals vs. Fitted Values')
axes[0].set_xlabel('Fitted mpg')
axes[0].set_ylabel('Residual')
stats.probplot(residuals, dist='norm', plot=axes[1])
axes[1].set_title('Q-Q Plot of Real Residuals')
plt.tight_layout()
plt.show()
No description has been provided for this image
shapiro_stat, shapiro_p = stats.shapiro(residuals)
print(f'Real Shapiro-Wilk test on residuals: statistic={shapiro_stat:.4f}, p-value={shapiro_p:.4f}')
Real Shapiro-Wilk test on residuals: statistic=0.9451, p-value=0.1044

Part 5: Two Cautions - Causation and Extrapolation#

print(f'Real weight range in this data: {x.min():.2f} to {x.max():.2f} (thousands of lbs)')
x_extreme = 8.0
pred_extreme = result.intercept + result.slope * x_extreme
print(f'Extrapolated prediction at {x_extreme} thousand lbs: {pred_extreme:.2f} mpg')
Real weight range in this data: 1.51 to 5.42 (thousands of lbs)
Extrapolated prediction at 8.0 thousand lbs: -5.47 mpg

And on causation: this real regression shows heavier cars genuinely tend to get worse real mileage, but it doesn't prove weight causes lower mpg in isolation. Heavier real cars in this dataset also tend to have bigger real engines and more real horsepower, and any of those could be doing some or all of the causal work. Untangling that requires more than one predictor at a time, exactly what video twelve's multiple regression adds.

Wrap-Up: What You Learned#

  • Fitting a simple linear regression on real mtcars data, and interpreting the slope and intercept in real units.
  • R-squared as the proportion of real variance explained, not a measure of correctness or causation.
  • Prediction intervals, quantifying genuine real uncertainty around a single new prediction.
  • Residual diagnostics: checking the real fitted-vs-residual plot and a Q-Q plot before trusting the model's other numbers.
  • Two real cautions: extrapolating beyond the observed real data range produces nonsense, and a strong real fit alone never proves causation.
  • Video twelve extends all of this to multiple predictors at once, including multicollinearity and a full statsmodels regression summary. Subscribe so it lands automatically see you there.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.