Mathew K Analytics

Lesson 26 · Probability and Statistics in python

Practical Regression Project: Predicting Housing Prices with Python and Data Analysis

In this lesson, you will learn how to predict housing prices using data and simple statistics. We will start from basic ideas, build up your skills, and…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Welcome to Hands-On Regression with Python!#

In this lesson, you will learn how to predict housing prices using data and simple statistics.

We will start from basic ideas, build up your skills, and work through a fun mini-project together.

import warnings; warnings.filterwarnings("ignore")
# Let us load libraries for data and math
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sns

What is Regression?#

Regression lets us predict a number from data. For example, we can estimate a house's price from its features.

You see regression in action in real life whenever a website suggests home values!

# Data setup
# We will use the popular California housing dataset from sklearn
from sklearn.datasets import fetch_california_housing
housing = fetch_california_housing(as_frame=True)
df = housing.frame
print("Shape:", df.shape)
df.head()
Shape: (20640, 9)
MedInc HouseAge AveRooms AveBedrms Population AveOccup Latitude Longitude MedHouseVal
0 8.3252 41.0 6.984127 1.023810 322.0 2.555556 37.88 -122.23 4.526
1 8.3014 21.0 6.238137 0.971880 2401.0 2.109842 37.86 -122.22 3.585
2 7.2574 52.0 8.288136 1.073446 496.0 2.802260 37.85 -122.24 3.521
3 5.6431 52.0 5.817352 1.073059 558.0 2.547945 37.85 -122.25 3.413
4 3.8462 52.0 6.281853 1.081081 565.0 2.181467 37.85 -122.25 3.422
# Let us see what features we have
print(df.columns.tolist())
# Look at some statistics for each column
df.describe()
['MedInc', 'HouseAge', 'AveRooms', 'AveBedrms', 'Population', 'AveOccup', 'Latitude', 'Longitude', 'MedHouseVal']
MedInc HouseAge AveRooms AveBedrms Population AveOccup Latitude Longitude MedHouseVal
count 20640.000000 20640.000000 20640.000000 20640.000000 20640.000000 20640.000000 20640.000000 20640.000000 20640.000000
mean 3.870671 28.639486 5.429000 1.096675 1425.476744 3.070655 35.631861 -119.569704 2.068558
std 1.899822 12.585558 2.474173 0.473911 1132.462122 10.386050 2.135952 2.003532 1.153956
min 0.499900 1.000000 0.846154 0.333333 3.000000 0.692308 32.540000 -124.350000 0.149990
25% 2.563400 18.000000 4.440716 1.006079 787.000000 2.429741 33.930000 -121.800000 1.196000
50% 3.534800 29.000000 5.229129 1.048780 1166.000000 2.818116 34.260000 -118.490000 1.797000
75% 4.743250 37.000000 6.052381 1.099526 1725.000000 3.282261 37.710000 -118.010000 2.647250
max 15.000100 52.000000 141.909091 34.066667 35682.000000 1243.333333 41.950000 -114.310000 5.000010

What do the columns mean?#

  • MedInc: Median income in the area (in tens of thousands)
  • HouseAge: Median age of houses
  • AveRooms: Average rooms per house
  • AveBedrms: Average bedrooms per house
  • Population: People living in the area
  • AveOccup: Average people in a house
  • Latitude and Longitude: Location
  • MedHouseVal: Median house value (what we want to predict!)
# Visualize the target: housing value distribution
plt.figure(figsize=(7,4))
sns.histplot(df['MedHouseVal'], bins=30, kde=True)
plt.xlabel('Median House Value (x $100,000)')
plt.title('Distribution of California House Prices')
plt.show()
No description has been provided for this image
# Check for missing data
missing = df.isnull().sum()
print(missing[missing > 0])
Series([], dtype: int64)

Correlation: Which features matter?#

Correlation tells us how closely two things move together. Strong correlation means knowing one value helps us guess the other.

# Calculate correlation with house value
corr = df.corr()
print(corr['MedHouseVal'].sort_values(ascending=False))
MedHouseVal    1.000000
MedInc         0.688075
AveRooms       0.151948
HouseAge       0.105623
AveOccup      -0.023737
Population    -0.024650
Longitude     -0.045967
AveBedrms     -0.046701
Latitude      -0.144160
Name: MedHouseVal, dtype: float64
# Scatter plot: Income vs. House Value
plt.figure(figsize=(7,4))
sns.scatterplot(x='MedInc', y='MedHouseVal', data=df, alpha=0.5)
plt.xlabel('Median Income')
plt.ylabel('Median House Value')
plt.title('Income vs Price')
plt.show()
No description has been provided for this image
# Split data into features (X) and target (y)
y = df['MedHouseVal']
X = df.drop('MedHouseVal', axis=1)
# For simplicity, let us use all features for now

What is Linear Regression?#

Linear regression fits a straight line to the data, so we can predict new values.

We are going to train a linear regression model to see what it learns.

# Split into train and test sets
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
print("Training rows:", X_train.shape[0], "| Test rows:", X_test.shape[0])
Training rows: 16512 | Test rows: 4128
# Build and train a linear regression model
from sklearn.linear_model import LinearRegression
model = LinearRegression()
model.fit(X_train, y_train)
print('Model trained.')
Model trained.
# Predict test values and evaluate the model
from sklearn.metrics import mean_squared_error, r2_score
y_pred = model.predict(X_test)
mse = mean_squared_error(y_test, y_pred)
r2 = r2_score(y_test, y_pred)
print("Test MSE:", round(mse, 3))
print("Test R^2:", round(r2, 3))
Test MSE: 0.556
Test R^2: 0.576
# Plot predicted vs actual house prices
plt.figure(figsize=(6,6))
plt.scatter(y_test, y_pred, alpha=0.4)
plt.xlabel('Actual Price')
plt.ylabel('Predicted Price')
plt.title('Predicted vs. Actual House Prices')
plt.plot([y_test.min(), y_test.max()], [y_test.min(), y_test.max()], 'r--')
plt.show()
No description has been provided for this image
# Show feature importances (model coefficients)
coef = pd.Series(model.coef_, index=X.columns)
coef = coef.sort_values()
plt.figure(figsize=(8, 5))
coef.plot(kind='barh')
plt.title('Feature Impact on House Value')
plt.xlabel('Coefficient weight')
plt.ylabel('Feature')
plt.show()
No description has been provided for this image

Beyond Linear Regression#

Other models, like decision trees, can find more complex links between features.

Let us try switching to a more powerful regression model!

# Try a Random Forest regressor
from sklearn.ensemble import RandomForestRegressor
rf = RandomForestRegressor(n_estimators=100, random_state=42)
rf.fit(X_train, y_train)
y_rf_pred = rf.predict(X_test)
mse_rf = mean_squared_error(y_test, y_rf_pred)
r2_rf = r2_score(y_test, y_rf_pred)
print("Random Forest Test MSE:", round(mse_rf,3))
print("Random Forest R^2:", round(r2_rf,3))
Random Forest Test MSE: 0.255
Random Forest R^2: 0.805
# Let us make a prediction for a new house
user_median_income = input("Enter median income for the area (in tens of thousands): ")
user_age = input("Enter median house age: ")
user_rooms = input("Enter average rooms: ")
user_bedrooms = input("Enter average bedrooms: ")
user_population = input("Enter population: ")
user_occup = input("Enter average occupation: ")
user_lat = input("Enter latitude: ")
user_long = input("Enter longitude: ")
new_house = np.array([[float(user_median_income), float(user_age), float(user_rooms), float(user_bedrooms), float(user_population), float(user_occup), float(user_lat), float(user_long)]])
predicted_val = rf.predict(new_house)[0]
print("Predicted median house value: $", round(predicted_val * 100000, 2))
Predicted median house value: $ 146765.0

Recap and Congrats!#

You have loaded real data, explored it, built and evaluated models, and made predictions.

That is a huge leap!

Next, try using your own data or changing the features to see how smart your model can be.

Final step: Challenge yourself!#

Can you try a different regression model, or predict home prices using just one feature?

Share your results in the comments and tell us what you discover.

For more hands-on Python, like and subscribe!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.