Mathew K Analytics

Lesson 29 · Python Fundamentals

Understanding Simple and Multiple Linear Regression in Python for Bible Study Applications

Today we will learn how to predict values using simple and multiple linear regression in Python. We will start with basics, explore real data, and build our…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb
 

Welcome to Python Linear Regression!#

Today we will learn how to predict values using simple and multiple linear regression in Python.

We will start with basics, explore real data, and build our own prediction models.

No experience needed. Let us begin!

# First, let us be sure warnings do not distract us.
import warnings
warnings.filterwarnings("ignore")

print("Ready to learn!")
Ready to learn!

What is Linear Regression?#

Linear regression is a way to find a straight-line relationship between numbers.

It is used to predict something using one or more clues, like guessing house prices from size and location.

It is one of the easiest, most important tools in data science!

# Let us try some simple math in Python. 
x = 10
y = 3 * x + 7
print("When x =", x, "then y =", y)
When x = 10 then y = 37

Loading Real Data: California Housing#

Let us use real-world data to make our learning interesting!

We will use the California Housing dataset.

It has data like home prices, number of rooms, and population, perfect for regression.

# Data setup
from sklearn.datasets import fetch_california_housing
import pandas as pd
california = fetch_california_housing(as_frame=True)
df = california.frame
print("Shape of data:", df.shape)
df.head()
Shape of data: (20640, 9)
MedInc HouseAge AveRooms AveBedrms Population AveOccup Latitude Longitude MedHouseVal
0 8.3252 41.0 6.984127 1.023810 322.0 2.555556 37.88 -122.23 4.526
1 8.3014 21.0 6.238137 0.971880 2401.0 2.109842 37.86 -122.22 3.585
2 7.2574 52.0 8.288136 1.073446 496.0 2.802260 37.85 -122.24 3.521
3 5.6431 52.0 5.817352 1.073059 558.0 2.547945 37.85 -122.25 3.413
4 3.8462 52.0 6.281853 1.081081 565.0 2.181467 37.85 -122.25 3.422
# Let us look at the column names. 
df.columns
Index(['MedInc', 'HouseAge', 'AveRooms', 'AveBedrms', 'Population', 'AveOccup',
       'Latitude', 'Longitude', 'MedHouseVal'],
      dtype='object')
# Let us see basic statistics. 
df.describe()
MedInc HouseAge AveRooms AveBedrms Population AveOccup Latitude Longitude MedHouseVal
count 20640.000000 20640.000000 20640.000000 20640.000000 20640.000000 20640.000000 20640.000000 20640.000000 20640.000000
mean 3.870671 28.639486 5.429000 1.096675 1425.476744 3.070655 35.631861 -119.569704 2.068558
std 1.899822 12.585558 2.474173 0.473911 1132.462122 10.386050 2.135952 2.003532 1.153956
min 0.499900 1.000000 0.846154 0.333333 3.000000 0.692308 32.540000 -124.350000 0.149990
25% 2.563400 18.000000 4.440716 1.006079 787.000000 2.429741 33.930000 -121.800000 1.196000
50% 3.534800 29.000000 5.229129 1.048780 1166.000000 2.818116 34.260000 -118.490000 1.797000
75% 4.743250 37.000000 6.052381 1.099526 1725.000000 3.282261 37.710000 -118.010000 2.647250
max 15.000100 52.000000 141.909091 34.066667 35682.000000 1243.333333 41.950000 -114.310000 5.000010
# What are we predicting? The target: MedHouseVal.
df['MedHouseVal'].head()
0    4.526
1    3.585
2    3.521
3    3.413
4    3.422
Name: MedHouseVal, dtype: float64

What makes a good input (feature)?#

Features are clues we use to make a prediction.

We want features that matter, like income or number of rooms.

Choosing good features helps make better predictions!

# Let us try simple linear regression with just one feature.
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LinearRegression

feature = 'MedInc'  # Median income
X = df[[feature]]
y = df['MedHouseVal']

X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

model = LinearRegression()
model.fit(X_train, y_train)

print("Model is trained!")
Model is trained!
# Let us predict home values for test data.
y_pred = model.predict(X_test)
print("First 5 Predictions:", y_pred[:5])
print("First 5 Actual Values:", list(y_test[:5]))
First 5 Predictions: [1.14958917 1.50606882 1.90393718 2.85059383 2.00663318]
First 5 Actual Values: [0.477, 0.458, 5.00001, 2.186, 2.78]
# Calculate how well our model did.
from sklearn.metrics import mean_squared_error
mse = mean_squared_error(y_test, y_pred)
print("Mean Squared Error:", mse)
Mean Squared Error: 0.7091157771765548
# What does our line look like? Let us plot it.
import matplotlib.pyplot as plt

plt.scatter(X_test, y_test, color="blue", label="Actual")
plt.plot(X_test, y_pred, color="red", label="Prediction")
plt.xlabel("Median Income")
plt.ylabel("Median House Value")
plt.title("Home Value Prediction - Simple Regression")
plt.legend()
plt.show()
No description has been provided for this image

Multiple Linear Regression#

Now let us use more than one feature, for better predictions.

This lets us capture more patterns in the data.

It is like using extra clues to guess the answer.

# Let us pick a few features.
features = ['MedInc', 'AveRooms', 'AveOccup']
X = df[features]

X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

multi_model = LinearRegression()
multi_model.fit(X_train, y_train)

multi_pred = multi_model.predict(X_test)

print("First 5 Predictions:", multi_pred[:5])
First 5 Predictions: [1.15826883 1.4999938  1.96250199 2.85239242 2.0039121 ]
# Let us measure this model's performance.
multi_mse = mean_squared_error(y_test, multi_pred)
print("Multi-feature MSE:", multi_mse)
print("Root MSE:", multi_mse ** 0.5)
Multi-feature MSE: 0.7006855912225248
Root MSE: 0.8370696453835397
# View feature importance (the weights).
importance = dict(zip(features, multi_model.coef_))
print("Feature Weights:")
for feat, coef in importance.items():
    print(feat, ":", coef)
    
Feature Weights:
MedInc : 0.43688279234997013
AveRooms : -0.04042961959300062
AveOccup : -0.003826017564824924
# Predict for your own values!
example = [[5.0, 6.0, 3.0]]
prediction = multi_model.predict(example)
print("Predicted home value:", prediction[0])
Predicted home value: 2.5384638846366294
# Try your own prediction using input().
income = float(input("Enter median income: "))
rooms = float(input("Enter average rooms: "))
occup = float(input("Enter average occupancy: "))
user_pred = multi_model.predict([[income, rooms, occup]])
print("Predicted house value:", user_pred[0])
Predicted house value: 1.9151007900295922
 

Common Problems and Best Practices#

Regression works best when relationships are fairly straight.

Too many features can confuse the model or cause overfitting.

Splitting data with random_state lets everyone repeat your work.

Always check results to be sure predictions make sense!

# Troubleshooting: What if you get a funny error?
try:
    broken = multi_model.predict([[None, 6.0, 3.0]])
    print(broken)
except Exception as e:
    print("Oops! Prediction failed. Reason:", e)
    
Oops! Prediction failed. Reason: unsupported operand type(s) for *: 'NoneType' and 'float'

Challenge: Try It Yourself!#

  1. Train a simple regression using 'AveRooms' as your only feature.

  2. Predict home values and plot predictions vs real values.

  3. Measure the mean squared error.

See if you can improve! Share your results below the video.

Recap#

Today you learned the basics of simple and multiple linear regression.

We explored real housing data, built our own models, made predictions, and found out how to improve them.

Practice makes perfect. Try again with different features!

Thank You and Next Steps#

Subscribe to the channel for more Python and data science tutorials.

Comment with what you tried and what you want to see next.

Practice, experiment, and you will be a data scientist in no time!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.