Lesson 26 · Probability and Statistics in python
Practical Regression Project: Predicting Housing Prices with Python and Data Analysis
In this lesson, you will learn how to predict housing prices using data and simple statistics. We will start from basic ideas, build up your skills, and…
- CourseProbability and Statistics in python
- Lesson26 of 35
- Video10 min
- FormatJupyter notebook · 15 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbWelcome to Hands-On Regression with Python!#
In this lesson, you will learn how to predict housing prices using data and simple statistics.
We will start from basic ideas, build up your skills, and work through a fun mini-project together.
import warnings; warnings.filterwarnings("ignore")
# Let us load libraries for data and math
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sns
What is Regression?#
Regression lets us predict a number from data. For example, we can estimate a house's price from its features.
You see regression in action in real life whenever a website suggests home values!
# Data setup
# We will use the popular California housing dataset from sklearn
from sklearn.datasets import fetch_california_housing
housing = fetch_california_housing(as_frame=True)
df = housing.frame
print("Shape:", df.shape)
df.head()
# Let us see what features we have
print(df.columns.tolist())
# Look at some statistics for each column
df.describe()
What do the columns mean?#
- MedInc: Median income in the area (in tens of thousands)
- HouseAge: Median age of houses
- AveRooms: Average rooms per house
- AveBedrms: Average bedrooms per house
- Population: People living in the area
- AveOccup: Average people in a house
- Latitude and Longitude: Location
- MedHouseVal: Median house value (what we want to predict!)
# Visualize the target: housing value distribution
plt.figure(figsize=(7,4))
sns.histplot(df['MedHouseVal'], bins=30, kde=True)
plt.xlabel('Median House Value (x $100,000)')
plt.title('Distribution of California House Prices')
plt.show()
# Check for missing data
missing = df.isnull().sum()
print(missing[missing > 0])
Correlation: Which features matter?#
Correlation tells us how closely two things move together. Strong correlation means knowing one value helps us guess the other.
# Calculate correlation with house value
corr = df.corr()
print(corr['MedHouseVal'].sort_values(ascending=False))
# Scatter plot: Income vs. House Value
plt.figure(figsize=(7,4))
sns.scatterplot(x='MedInc', y='MedHouseVal', data=df, alpha=0.5)
plt.xlabel('Median Income')
plt.ylabel('Median House Value')
plt.title('Income vs Price')
plt.show()
# Split data into features (X) and target (y)
y = df['MedHouseVal']
X = df.drop('MedHouseVal', axis=1)
# For simplicity, let us use all features for now
What is Linear Regression?#
Linear regression fits a straight line to the data, so we can predict new values.
We are going to train a linear regression model to see what it learns.
# Split into train and test sets
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
print("Training rows:", X_train.shape[0], "| Test rows:", X_test.shape[0])
# Build and train a linear regression model
from sklearn.linear_model import LinearRegression
model = LinearRegression()
model.fit(X_train, y_train)
print('Model trained.')
# Predict test values and evaluate the model
from sklearn.metrics import mean_squared_error, r2_score
y_pred = model.predict(X_test)
mse = mean_squared_error(y_test, y_pred)
r2 = r2_score(y_test, y_pred)
print("Test MSE:", round(mse, 3))
print("Test R^2:", round(r2, 3))
# Plot predicted vs actual house prices
plt.figure(figsize=(6,6))
plt.scatter(y_test, y_pred, alpha=0.4)
plt.xlabel('Actual Price')
plt.ylabel('Predicted Price')
plt.title('Predicted vs. Actual House Prices')
plt.plot([y_test.min(), y_test.max()], [y_test.min(), y_test.max()], 'r--')
plt.show()
# Show feature importances (model coefficients)
coef = pd.Series(model.coef_, index=X.columns)
coef = coef.sort_values()
plt.figure(figsize=(8, 5))
coef.plot(kind='barh')
plt.title('Feature Impact on House Value')
plt.xlabel('Coefficient weight')
plt.ylabel('Feature')
plt.show()
Beyond Linear Regression#
Other models, like decision trees, can find more complex links between features.
Let us try switching to a more powerful regression model!
# Try a Random Forest regressor
from sklearn.ensemble import RandomForestRegressor
rf = RandomForestRegressor(n_estimators=100, random_state=42)
rf.fit(X_train, y_train)
y_rf_pred = rf.predict(X_test)
mse_rf = mean_squared_error(y_test, y_rf_pred)
r2_rf = r2_score(y_test, y_rf_pred)
print("Random Forest Test MSE:", round(mse_rf,3))
print("Random Forest R^2:", round(r2_rf,3))
# Let us make a prediction for a new house
user_median_income = input("Enter median income for the area (in tens of thousands): ")
user_age = input("Enter median house age: ")
user_rooms = input("Enter average rooms: ")
user_bedrooms = input("Enter average bedrooms: ")
user_population = input("Enter population: ")
user_occup = input("Enter average occupation: ")
user_lat = input("Enter latitude: ")
user_long = input("Enter longitude: ")
new_house = np.array([[float(user_median_income), float(user_age), float(user_rooms), float(user_bedrooms), float(user_population), float(user_occup), float(user_lat), float(user_long)]])
predicted_val = rf.predict(new_house)[0]
print("Predicted median house value: $", round(predicted_val * 100000, 2))
Recap and Congrats!#
You have loaded real data, explored it, built and evaluated models, and made predictions.
That is a huge leap!
Next, try using your own data or changing the features to see how smart your model can be.
Final step: Challenge yourself!#
Can you try a different regression model, or predict home prices using just one feature?
Share your results in the comments and tell us what you discover.
For more hands-on Python, like and subscribe!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



