Lesson 29 · Python Fundamentals
Understanding Simple and Multiple Linear Regression in Python for Bible Study Applications
Today we will learn how to predict values using simple and multiple linear regression in Python. We will start with basics, explore real data, and build our…
- CoursePython Fundamentals
- Lesson29 of 22
- Video12 min
- FormatJupyter notebook · 18 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynb
Welcome to Python Linear Regression!#
Today we will learn how to predict values using simple and multiple linear regression in Python.
We will start with basics, explore real data, and build our own prediction models.
No experience needed. Let us begin!
# First, let us be sure warnings do not distract us.
import warnings
warnings.filterwarnings("ignore")
print("Ready to learn!")
What is Linear Regression?#
Linear regression is a way to find a straight-line relationship between numbers.
It is used to predict something using one or more clues, like guessing house prices from size and location.
It is one of the easiest, most important tools in data science!
# Let us try some simple math in Python.
x = 10
y = 3 * x + 7
print("When x =", x, "then y =", y)
Loading Real Data: California Housing#
Let us use real-world data to make our learning interesting!
We will use the California Housing dataset.
It has data like home prices, number of rooms, and population, perfect for regression.
# Data setup
from sklearn.datasets import fetch_california_housing
import pandas as pd
california = fetch_california_housing(as_frame=True)
df = california.frame
print("Shape of data:", df.shape)
df.head()
# Let us look at the column names.
df.columns
# Let us see basic statistics.
df.describe()
# What are we predicting? The target: MedHouseVal.
df['MedHouseVal'].head()
What makes a good input (feature)?#
Features are clues we use to make a prediction.
We want features that matter, like income or number of rooms.
Choosing good features helps make better predictions!
# Let us try simple linear regression with just one feature.
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LinearRegression
feature = 'MedInc' # Median income
X = df[[feature]]
y = df['MedHouseVal']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
model = LinearRegression()
model.fit(X_train, y_train)
print("Model is trained!")
# Let us predict home values for test data.
y_pred = model.predict(X_test)
print("First 5 Predictions:", y_pred[:5])
print("First 5 Actual Values:", list(y_test[:5]))
# Calculate how well our model did.
from sklearn.metrics import mean_squared_error
mse = mean_squared_error(y_test, y_pred)
print("Mean Squared Error:", mse)
# What does our line look like? Let us plot it.
import matplotlib.pyplot as plt
plt.scatter(X_test, y_test, color="blue", label="Actual")
plt.plot(X_test, y_pred, color="red", label="Prediction")
plt.xlabel("Median Income")
plt.ylabel("Median House Value")
plt.title("Home Value Prediction - Simple Regression")
plt.legend()
plt.show()
Multiple Linear Regression#
Now let us use more than one feature, for better predictions.
This lets us capture more patterns in the data.
It is like using extra clues to guess the answer.
# Let us pick a few features.
features = ['MedInc', 'AveRooms', 'AveOccup']
X = df[features]
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
multi_model = LinearRegression()
multi_model.fit(X_train, y_train)
multi_pred = multi_model.predict(X_test)
print("First 5 Predictions:", multi_pred[:5])
# Let us measure this model's performance.
multi_mse = mean_squared_error(y_test, multi_pred)
print("Multi-feature MSE:", multi_mse)
print("Root MSE:", multi_mse ** 0.5)
# View feature importance (the weights).
importance = dict(zip(features, multi_model.coef_))
print("Feature Weights:")
for feat, coef in importance.items():
print(feat, ":", coef)
# Predict for your own values!
example = [[5.0, 6.0, 3.0]]
prediction = multi_model.predict(example)
print("Predicted home value:", prediction[0])
# Try your own prediction using input().
income = float(input("Enter median income: "))
rooms = float(input("Enter average rooms: "))
occup = float(input("Enter average occupancy: "))
user_pred = multi_model.predict([[income, rooms, occup]])
print("Predicted house value:", user_pred[0])
Common Problems and Best Practices#
Regression works best when relationships are fairly straight.
Too many features can confuse the model or cause overfitting.
Splitting data with random_state lets everyone repeat your work.
Always check results to be sure predictions make sense!
# Troubleshooting: What if you get a funny error?
try:
broken = multi_model.predict([[None, 6.0, 3.0]])
print(broken)
except Exception as e:
print("Oops! Prediction failed. Reason:", e)
Challenge: Try It Yourself!#
Train a simple regression using 'AveRooms' as your only feature.
Predict home values and plot predictions vs real values.
Measure the mean squared error.
See if you can improve! Share your results below the video.
Recap#
Today you learned the basics of simple and multiple linear regression.
We explored real housing data, built our own models, made predictions, and found out how to improve them.
Practice makes perfect. Try again with different features!
Thank You and Next Steps#
Subscribe to the channel for more Python and data science tutorials.
Comment with what you tried and what you want to see next.
Practice, experiment, and you will be a data scientist in no time!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



