Mathew K Analytics

Lesson 77 · Data Science Projects

Predicting House Prices with Regression Using the Ames Housing Dataset

Welcome to this hands-on introduction to regression using real housing data. In this lesson, we will predict house prices with the Ames Housing Dataset. You…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Week 45: House Price Prediction with Ames Housing Data#

Welcome to this hands-on introduction to regression using real housing data.

In this lesson, we will predict house prices with the Ames Housing Dataset.

You will learn to load data, explore relationships, build your first regression model, and make predictions.

House price prediction is a core data mining project and a popular interview question.

# Import libraries and suppress warnings
import warnings; warnings.filterwarnings("ignore")
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns
np.random.seed(42)
# Data setup (Ames Housing Dataset)
url = 'https://raw.githubusercontent.com/ageron/handson-ml/master/datasets/housing/housing.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
(20640, 10)
   longitude  latitude  housing_median_age  total_rooms  total_bedrooms  \
0    -122.23     37.88                41.0        880.0           129.0   
1    -122.22     37.86                21.0       7099.0          1106.0   
2    -122.24     37.85                52.0       1467.0           190.0   

   population  households  median_income  median_house_value ocean_proximity  
0       322.0       126.0         8.3252            452600.0        NEAR BAY  
1      2401.0      1138.0         8.3014            358500.0        NEAR BAY  
2       496.0       177.0         7.2574            352100.0        NEAR BAY  

Why use regression for house prices?#

Regression helps us predict values, like a home's market price, based on its features: bedrooms, location, size, and more.

This is very useful for real estate, banks, and new home buyers.

# Basic info about the dataset
df.info()
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 20640 entries, 0 to 20639
Data columns (total 10 columns):
 #   Column              Non-Null Count  Dtype  
---  ------              --------------  -----  
 0   longitude           20640 non-null  float64
 1   latitude            20640 non-null  float64
 2   housing_median_age  20640 non-null  float64
 3   total_rooms         20640 non-null  float64
 4   total_bedrooms      20433 non-null  float64
 5   population          20640 non-null  float64
 6   households          20640 non-null  float64
 7   median_income       20640 non-null  float64
 8   median_house_value  20640 non-null  float64
 9   ocean_proximity     20640 non-null  object 
dtypes: float64(9), object(1)
memory usage: 1.6+ MB
# Check for missing values in the data
print(df.isnull().sum())
longitude               0
latitude                0
housing_median_age      0
total_rooms             0
total_bedrooms        207
population              0
households              0
median_income           0
median_house_value      0
ocean_proximity         0
dtype: int64
# Drop rows with missing values, for simplicity in this lesson
df_clean = df.dropna()
print('Shape after cleaning:', df_clean.shape)
Shape after cleaning: (20433, 10)
# Show basic statistics of numeric columns
print(df_clean.describe().T)
                      count           mean            std         min  \
longitude           20433.0    -119.570689       2.003578   -124.3500   
latitude            20433.0      35.633221       2.136348     32.5400   
housing_median_age  20433.0      28.633094      12.591805      1.0000   
total_rooms         20433.0    2636.504233    2185.269567      2.0000   
total_bedrooms      20433.0     537.870553     421.385070      1.0000   
population          20433.0    1424.946949    1133.208490      3.0000   
households          20433.0     499.433465     382.299226      1.0000   
median_income       20433.0       3.871162       1.899291      0.4999   
median_house_value  20433.0  206864.413155  115435.667099  14999.0000   

                            25%          50%         75%          max  
longitude             -121.8000    -118.4900    -118.010    -114.3100  
latitude                33.9300      34.2600      37.720      41.9500  
housing_median_age      18.0000      29.0000      37.000      52.0000  
total_rooms           1450.0000    2127.0000    3143.000   39320.0000  
total_bedrooms         296.0000     435.0000     647.000    6445.0000  
population             787.0000    1166.0000    1722.000   35682.0000  
households             280.0000     409.0000     604.000    6082.0000  
median_income            2.5637       3.5365       4.744      15.0001  
median_house_value  119500.0000  179700.0000  264700.000  500001.0000  
# Visualize the distribution of house prices
plt.figure(figsize=(8,4))
sns.histplot(df_clean['median_house_value'], kde=True, color='blue')
plt.title('Distribution of House Prices')
plt.xlabel('House Price (USD)')
plt.ylabel('Count')
plt.show()
No description has been provided for this image

Exploring relationships: Features vs Price#

Some factors impact price more than others. Let us investigate!

# Compare house price and median_income
plt.figure(figsize=(8,5))
sns.scatterplot(x='median_income', y='median_house_value', data=df_clean, alpha=0.3)
plt.title('Income vs House Price')
plt.xlabel('Median Income')
plt.ylabel('House Price (USD)')
plt.show()
No description has been provided for this image
# Select features and target variable for regression
X = df_clean[['median_income', 'housing_median_age', 'total_rooms', 'population']]
y = df_clean['median_house_value']
# Split into train and test sets for fair evaluation
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
# Build a simple linear regression model
from sklearn.linear_model import LinearRegression
reg = LinearRegression()
reg.fit(X_train, y_train)
LinearRegression()
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
# Predict house prices with our trained model
y_pred = reg.predict(X_test)
print('Predictions (first 5):', y_pred[:5])
Predictions (first 5): [156334.11022268 192526.60815285 184555.54953227 164617.27693086
 169212.14472774]
# Evaluate: how good is our model?
from sklearn.metrics import mean_absolute_error, mean_squared_error, r2_score
mae = mean_absolute_error(y_test, y_pred)
rmse = np.sqrt(mean_squared_error(y_test, y_pred))
r2 = r2_score(y_test, y_pred)
print('Mean Absolute Error:', mae)
print('Root Mean Squared Error:', rmse)
print('R^2 Score:', r2)
Mean Absolute Error: 60866.56099957577
Root Mean Squared Error: 81256.38115652399
R^2 Score: 0.5171837519405796
# Visualize actual vs predicted price for test set
plt.figure(figsize=(6,6))
plt.scatter(y_test, y_pred, alpha=0.4)
plt.xlabel('Actual Price')
plt.ylabel('Predicted Price')
plt.title('Actual vs Predicted House Prices')
plt.plot([y_test.min(), y_test.max()], [y_test.min(), y_test.max()], 'r--')
plt.show()
No description has been provided for this image

Practice: Try your own prediction#

Let us make our own house price prediction. Enter a few numbers when prompted.

# Predict price for your own custom data
income = float(input("Enter median_income (between 1 and 15): "))
age = float(input("Enter housing_median_age (between 1 and 50): "))
rooms = float(input("Enter total_rooms (e.g., 100-20000): "))
population = float(input("Enter population for that block (e.g., 100-36000): "))
user_X = np.array([[income, age, rooms, population]])
user_pred = reg.predict(user_X)[0]
print("Predicted house price: $", round(user_pred, 2))
Predicted house price: $ 312419.07

Challenge: Extend this notebook#

  1. Try adding more features to the model and check if error goes down.
  2. Try replacing LinearRegression with a RandomForestRegressor.
  3. Visualize how feature importances change.

Lesson recap and next steps#

You learned to:

  • Load and clean real housing data
  • Explore and plot data
  • Build, test, and evaluate a regression model
  • Make your own predictions

Keep practicing and try out new datasets to build your confidence!

Thank you! Practice and Subscribe#

That is all for house price regression.

Practice, try the challenge, and subscribe to stay on track with your learning journey!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.