Lesson 1 · ML algorithms deep dive
Linear Regression from Scratch in Python | ML Algorithms #1
The first video of a 12-part series: one real algorithm per video, covered in full depth. Real classic car data: fuel economy, weight, horsepower, and…
- CourseML algorithms deep dive
- Lesson1 of 12
- Video25 min
- FormatJupyter notebook · 17 code cells
- Data1 dataset
What you'll learn
Datasets used in this lesson
Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.
- mtcars.csv1.7 KB
📓 Full notebook
Download .ipynbML Algorithms Deep-Dive, Video 1: Linear Regression#
- The first video of a 12-part series: one real algorithm per video, covered in full depth.
- Real classic car data: fuel economy, weight, horsepower, and displacement.
- Let's get into it.
Before You Start#
- Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
- You'll need pandas, NumPy, Matplotlib, SciPy, scikit-learn, and statsmodels.
- Place mtcars.csv in the same folder as this notebook.
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
from scipy import stats
from sklearn.linear_model import LinearRegression, Ridge, Lasso
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import KFold, cross_val_score
from statsmodels.stats.outliers_influence import variance_inflation_factor
from statsmodels.stats.stattools import durbin_watson
import statsmodels.api as sm
cars = pd.read_csv('mtcars.csv')
print(f'Real cars in this dataset: {len(cars)}')
print(cars[['model', 'mpg', 'wt', 'hp', 'disp']].head())
Part 1: The Problem Linear Regression Solves#
plt.figure(figsize=(8, 5))
plt.scatter(cars['wt'], cars['mpg'], color='steelblue')
plt.title('Real Fuel Economy vs. Real Car Weight')
plt.xlabel('Weight (1000 lbs)')
plt.ylabel('Miles per Gallon')
plt.show()
print(f"Real correlation between weight and mpg: {cars['wt'].corr(cars['mpg']):.4f}")
Part 2: The Math Behind the Line#
X_simple = cars[['wt']].values
y = cars['mpg'].values
simple_model = LinearRegression().fit(X_simple, y)
print(f'Real intercept: {simple_model.intercept_:.4f}')
print(f'Real slope: {simple_model.coef_[0]:.4f}')
print(f'Real R-squared: {simple_model.score(X_simple, y):.4f}')
plt.figure(figsize=(8, 5))
plt.scatter(cars['wt'], cars['mpg'], color='steelblue', label='Real cars')
x_line = np.linspace(cars['wt'].min(), cars['wt'].max(), 50)
y_line = simple_model.predict(x_line.reshape(-1, 1))
plt.plot(x_line, y_line, color='darkred', label='Fitted line')
plt.title('Real Fuel Economy vs. Weight, with the Fitted Regression Line')
plt.xlabel('Weight (1000 lbs)')
plt.ylabel('Miles per Gallon')
plt.legend()
plt.show()
Part 3: Multiple Regression#
X_multi = cars[['wt', 'hp', 'disp']].values
multi_model = LinearRegression().fit(X_multi, y)
print(f'Real intercept: {multi_model.intercept_:.4f}')
for name, coef in zip(['wt', 'hp', 'disp'], multi_model.coef_):
print(f'Real coefficient for {name}: {coef:.4f}')
print(f'Real R-squared: {multi_model.score(X_multi, y):.4f}')
Holding horsepower and displacement fixed, each additional real 1000 pounds still costs roughly 3.8 real miles per gallon, the dominant real effect. Horsepower's own real coefficient is small and displacement's is nearly zero, even though displacement alone correlated with mpg at negative 0.85. That mismatch is worth investigating before trusting these real numbers.
Part 4: Multicollinearity#
print(cars[['wt', 'hp', 'disp']].corr())
X_with_const = sm.add_constant(cars[['wt', 'hp', 'disp']])
vif_values = [variance_inflation_factor(X_with_const.values, i) for i in range(X_with_const.shape[1])]
for name, vif in zip(X_with_const.columns, vif_values):
print(f'Real VIF for {name}: {vif:.2f}')
This is the real explanation: displacement doesn't lack real predictive power on its own, its real information is largely redundant with weight once weight is already in the model, so the fitting process can't cleanly separate their individual real contributions. That's exactly why part three's coefficients shouldn't be read as 'displacement doesn't matter.'
Part 5: Checking the Assumptions#
fitted_values = multi_model.predict(X_multi)
residuals = y - fitted_values
plt.figure(figsize=(8, 5))
plt.scatter(fitted_values, residuals, color='steelblue')
plt.axhline(0, color='darkred', linestyle='--')
plt.title('Real Residuals vs. Real Fitted Values')
plt.xlabel('Fitted mpg')
plt.ylabel('Residual')
plt.show()
bp_stat, bp_p, _, _ = sm.stats.diagnostic.het_breuschpagan(residuals, X_with_const)
print(f'Real Breusch-Pagan statistic: {bp_stat:.4f}, p-value: {bp_p:.4f}')
jb_stat, jb_p = stats.jarque_bera(residuals)
print(f'Real Jarque-Bera statistic: {jb_stat:.4f}, p-value: {jb_p:.4f}')
sh_stat, sh_p = stats.shapiro(residuals)
print(f'Real Shapiro-Wilk statistic: {sh_stat:.4f}, p-value: {sh_p:.4f}')
dw_stat = durbin_watson(residuals)
print(f'Real Durbin-Watson statistic: {dw_stat:.4f}')
Two real normality tests disagreeing at a small sample size isn't a bug, it's an honest, common real outcome: Shapiro-Wilk has more power to detect small real deviations, Jarque-Bera less so. With only thirty-two real cars, don't treat either verdict as final; look at the real residual plot itself, which shows nothing dramatically non-normal.
Part 6: Why a Single Train/Test Split Isn't Enough Here#
kfold = KFold(n_splits=5, shuffle=True, random_state=42)
cv_r2 = cross_val_score(LinearRegression(), X_multi, y, cv=kfold, scoring='r2')
cv_mae = -cross_val_score(LinearRegression(), X_multi, y, cv=kfold, scoring='neg_mean_absolute_error')
print(f'Real per-fold R-squared: {[round(v, 4) for v in cv_r2]}')
print(f'Real mean CV R-squared: {cv_r2.mean():.4f}')
print(f'Real per-fold MAE: {[round(v, 4) for v in cv_mae]}')
print(f'Real mean CV MAE: {cv_mae.mean():.4f}')
The real spread across folds tells its own story: one real fold scores an R-squared as low as 0.44, another as high as 0.87, purely from which six or seven real cars happened to land in the held-out group. The full-data R-squared of 0.83 from part three was optimistic; the honest real out-of-sample estimate is closer to 0.74, with real uncertainty around it that a single number can't convey.
Part 7: Regularization: Ridge and Lasso#
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X_multi)
ols_scaled = LinearRegression().fit(X_scaled, y)
print(f'Real standardized OLS coefficients: {[round(c, 4) for c in ols_scaled.coef_]}')
for alpha in [0.1, 1.0, 10.0]:
ridge = Ridge(alpha=alpha).fit(X_scaled, y)
lasso = Lasso(alpha=alpha).fit(X_scaled, y)
print(f'alpha={alpha}: Ridge {[round(c, 4) for c in ridge.coef_]}, Lasso {[round(c, 4) for c in lasso.coef_]}')
for alpha in [0.1, 1.0, 10.0]:
ridge_cv = cross_val_score(Ridge(alpha=alpha), X_scaled, y, cv=kfold, scoring='r2').mean()
lasso_cv = cross_val_score(Lasso(alpha=alpha), X_scaled, y, cv=kfold, scoring='r2').mean()
print(f'alpha={alpha}: Real CV R-squared Ridge={ridge_cv:.4f}, Lasso={lasso_cv:.4f}')
ols_cv_scaled = cross_val_score(LinearRegression(), X_scaled, y, cv=kfold, scoring='r2').mean()
print(f'Real CV R-squared, plain OLS: {ols_cv_scaled:.4f}')
A real, honest result: at alpha equals 10, Ridge's cross-validated R-squared actually rises to about 0.77, mildly beating plain OLS's 0.74, because shrinking the collinear coefficients reduces real overfitting to noise. Lasso at that same strength collapses to negative territory, worse than just predicting the real average mpg, because it zeroed out every real coefficient entirely. Regularization isn't automatically better, and more of it isn't automatically better either; it has to be tuned and checked against a real held-out score, not assumed.
Part 8: Evaluation Metrics, Recapped#
n_obs = len(y)
n_predictors = X_multi.shape[1]
r_squared = multi_model.score(X_multi, y)
adjusted_r_squared = 1 - (1 - r_squared) * (n_obs - 1) / (n_obs - n_predictors - 1)
print(f'Real R-squared: {r_squared:.4f}')
print(f'Real adjusted R-squared: {adjusted_r_squared:.4f}')
Mean absolute error is the real average size of a miss, in real miles per gallon, regardless of direction. Root mean squared error is similar but squares errors first, so a few real large misses pull it up more than MAE. R-squared and adjusted R-squared describe real explanatory power; MAE and RMSE describe real prediction error in the target's own real units, which is usually the more useful number to report to a non-technical real audience.
Part 9: Strengths, Weaknesses, and When to Use It#
- Strength: every coefficient has a direct, real, interpretable meaning.
- Strength: fast to fit, no hyperparameter search required for the plain version.
- Weakness: assumes a real linear relationship; real curved patterns need transformed features or a different algorithm entirely.
- Weakness: sensitive to real multicollinearity and to real outliers, as parts four and five showed directly.
- Use it when interpretability matters as much as accuracy, and when the real relationship is plausibly close to linear.
Part 10: Saving Your Work#
final_model = Ridge(alpha=10).fit(X_scaled, y)
predicted_mpg = final_model.predict(X_scaled)
plt.figure(figsize=(8, 5))
plt.scatter(y, predicted_mpg, color='steelblue')
plt.plot([y.min(), y.max()], [y.min(), y.max()], color='darkred', linestyle='--')
plt.title('Real Actual vs. Predicted MPG (Final Ridge Model)')
plt.xlabel('Actual MPG')
plt.ylabel('Predicted MPG')
plt.savefig('linear_regression_final_fit.png', dpi=150)
plt.show()
import joblib
joblib.dump(final_model, 'linear_regression_model.joblib')
reloaded_model = joblib.load('linear_regression_model.joblib')
print(f'Real original model R-squared: {final_model.score(X_scaled, y):.4f}')
print(f'Real reloaded model R-squared: {reloaded_model.score(X_scaled, y):.4f}')
Wrap-Up: What You Learned#
- The real math behind a fitted regression line, and how it extends to multiple real predictors.
- Diagnosing real multicollinearity with correlation and VIF, and why it distorts coefficients rather than predictions.
- Testing every real OLS assumption directly: linearity, constant variance, normality, and independence.
- Why a real small dataset needs cross-validation, not a single lucky or unlucky train/test split.
- Ridge and Lasso regularization, including a real case where mild Ridge helps and heavy Lasso actively hurts.
- Video two moves to classification: logistic regression, on a real, larger medical dataset. Subscribe so it lands automatically see you there.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



