Lesson 77 · Data Science Projects
Predicting House Prices with Regression Using the Ames Housing Dataset
Welcome to this hands-on introduction to regression using real housing data. In this lesson, we will predict house prices with the Ames Housing Dataset. You…
- CourseData Science Projects
- Lesson77 of 33
- Video19 min
- FormatJupyter notebook · 15 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbWeek 45: House Price Prediction with Ames Housing Data#
Welcome to this hands-on introduction to regression using real housing data.
In this lesson, we will predict house prices with the Ames Housing Dataset.
You will learn to load data, explore relationships, build your first regression model, and make predictions.
House price prediction is a core data mining project and a popular interview question.
# Import libraries and suppress warnings
import warnings; warnings.filterwarnings("ignore")
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns
np.random.seed(42)
# Data setup (Ames Housing Dataset)
url = 'https://raw.githubusercontent.com/ageron/handson-ml/master/datasets/housing/housing.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
Why use regression for house prices?#
Regression helps us predict values, like a home's market price, based on its features: bedrooms, location, size, and more.
This is very useful for real estate, banks, and new home buyers.
# Basic info about the dataset
df.info()
# Check for missing values in the data
print(df.isnull().sum())
# Drop rows with missing values, for simplicity in this lesson
df_clean = df.dropna()
print('Shape after cleaning:', df_clean.shape)
# Show basic statistics of numeric columns
print(df_clean.describe().T)
# Visualize the distribution of house prices
plt.figure(figsize=(8,4))
sns.histplot(df_clean['median_house_value'], kde=True, color='blue')
plt.title('Distribution of House Prices')
plt.xlabel('House Price (USD)')
plt.ylabel('Count')
plt.show()
Exploring relationships: Features vs Price#
Some factors impact price more than others. Let us investigate!
# Compare house price and median_income
plt.figure(figsize=(8,5))
sns.scatterplot(x='median_income', y='median_house_value', data=df_clean, alpha=0.3)
plt.title('Income vs House Price')
plt.xlabel('Median Income')
plt.ylabel('House Price (USD)')
plt.show()
# Select features and target variable for regression
X = df_clean[['median_income', 'housing_median_age', 'total_rooms', 'population']]
y = df_clean['median_house_value']
# Split into train and test sets for fair evaluation
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
# Build a simple linear regression model
from sklearn.linear_model import LinearRegression
reg = LinearRegression()
reg.fit(X_train, y_train)
# Predict house prices with our trained model
y_pred = reg.predict(X_test)
print('Predictions (first 5):', y_pred[:5])
# Evaluate: how good is our model?
from sklearn.metrics import mean_absolute_error, mean_squared_error, r2_score
mae = mean_absolute_error(y_test, y_pred)
rmse = np.sqrt(mean_squared_error(y_test, y_pred))
r2 = r2_score(y_test, y_pred)
print('Mean Absolute Error:', mae)
print('Root Mean Squared Error:', rmse)
print('R^2 Score:', r2)
# Visualize actual vs predicted price for test set
plt.figure(figsize=(6,6))
plt.scatter(y_test, y_pred, alpha=0.4)
plt.xlabel('Actual Price')
plt.ylabel('Predicted Price')
plt.title('Actual vs Predicted House Prices')
plt.plot([y_test.min(), y_test.max()], [y_test.min(), y_test.max()], 'r--')
plt.show()
Practice: Try your own prediction#
Let us make our own house price prediction. Enter a few numbers when prompted.
# Predict price for your own custom data
income = float(input("Enter median_income (between 1 and 15): "))
age = float(input("Enter housing_median_age (between 1 and 50): "))
rooms = float(input("Enter total_rooms (e.g., 100-20000): "))
population = float(input("Enter population for that block (e.g., 100-36000): "))
user_X = np.array([[income, age, rooms, population]])
user_pred = reg.predict(user_X)[0]
print("Predicted house price: $", round(user_pred, 2))
Challenge: Extend this notebook#
- Try adding more features to the model and check if error goes down.
- Try replacing LinearRegression with a RandomForestRegressor.
- Visualize how feature importances change.
Lesson recap and next steps#
You learned to:
- Load and clean real housing data
- Explore and plot data
- Build, test, and evaluate a regression model
- Make your own predictions
Keep practicing and try out new datasets to build your confidence!
Thank you! Practice and Subscribe#
That is all for house price regression.
Practice, try the challenge, and subscribe to stay on track with your learning journey!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



