Lesson 30 · Data Science Projects
Building a House Price Prediction Model with Python: Data Science and Machine Learning
In this lesson, you will learn how to predict house prices using real world data. You will see the full process: loading data, inspecting it, doing…
- CourseData Science Projects
- Lesson30 of 33
- Video15 min
- FormatJupyter notebook · 16 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbHouse Price Prediction using Machine Learning#
In this lesson, you will learn how to predict house prices using real world data.
You will see the full process: loading data, inspecting it, doing analysis, cleaning, building a model, and making predictions.
House price prediction is a classic example in data mining and real estate analytics.
You will also learn to visualize, explain, and improve model results.
This is for absolute beginners. No previous coding or machine learning experience is needed.
Let us get started and discover the power of Python in data science!
Lesson Goals#
Understand the process of predicting house prices using data.
Explore and prepare real house price data with pandas.
Build and train a simple machine learning model.
Evaluate model accuracy and make predictions.
Learn to visualize data and results.
By the end, you will be able to apply this workflow to other datasets.
Remember to subscribe if you enjoy these beginner lessons!
# Suppress warnings to make outputs clean
import warnings; warnings.filterwarnings("ignore")
import numpy as np
np.random.seed(42)
Data setup#
- We are using the California Housing Dataset for this lesson.
- This data is real and describes housing in California districts.
- It includes features like number of rooms, population, and median value.
- We will load the data with pandas and preview its shape and contents.
# Load the California housing dataset
import pandas as pd
url = 'https://raw.githubusercontent.com/ageron/handson-ml/master/datasets/housing/housing.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
# Let us check the column names and types
print(df.dtypes)
# Summary statistics help us see data shape
print(df.describe())
Checking for missing data#
- Real world data is not always perfect.
- Let us check whether we have any missing values in our dataset.
- Missing values can cause problems when building models, so it is best to catch them early.
# Find missing values by column
print(df.isnull().sum())
# Remove rows with missing values
df = df.dropna()
print(df.shape)
Visualizing the target variable#
- The target variable is what we want to predict. In this dataset, it is 'median_house_value'.
- Let us look at its distribution to understand its range and spread.
- Visualizing the target helps spot any extreme or unusual values.
# Plot the distribution of median house values
import matplotlib.pyplot as plt
plt.figure(figsize=(8,4))
plt.hist(df['median_house_value'], bins=30, color='skyblue')
plt.title('Distribution of Median House Value')
plt.xlabel('Median House Value')
plt.ylabel('Number of Districts')
plt.show()
# Quick look at correlation with house value
corr = df.corr()
print(corr['median_house_value'].sort_values(ascending=False))
# Visualize longitude and latitude
plt.figure(figsize=(8,6))
plt.scatter(df['longitude'], df['latitude'], alpha=0.2,
c=df['median_house_value'], cmap='viridis')
plt.colorbar(label='Median House Value')
plt.xlabel('Longitude')
plt.ylabel('Latitude')
plt.title('House Prices by Location')
plt.show()
Feature selection#
To build a prediction model, we need to choose which features to use.
For this lesson, we will use simple numeric columns.
Picking the right features can make models perform better and be easier to understand.
# Select numeric features
features = [
'median_income',
'housing_median_age',
'total_rooms',
'total_bedrooms',
'population',
'households'
]
X = df[features]
y = df['median_house_value']
print(X.shape, y.shape)
# Split into training and test sets
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42)
print(X_train.shape, X_test.shape)
Building a linear regression model#
- A linear regression model is a simple and popular choice for predicting numbers.
- It learns to draw a straight line that best fits the training data.
- Let us build and train the model using scikit learn, a friendly Python package.
# Train a linear regression model
from sklearn.linear_model import LinearRegression
lr = LinearRegression()
lr.fit(X_train, y_train)
# Predict prices for test data
y_pred = lr.predict(X_test)
print(y_pred[:5])
# Measure model accuracy
from sklearn.metrics import mean_squared_error, r2_score
mse = mean_squared_error(y_test, y_pred)
rmse = mse ** 0.5
r2 = r2_score(y_test, y_pred)
print(f"Root Mean Squared Error: {rmse:.2f}")
print(f"R2 Score: {r2:.3f}")
# Plot predicted vs actual house prices
plt.figure(figsize=(6,6))
plt.scatter(y_test, y_pred, alpha=0.2, color='purple')
plt.xlabel('Actual Price')
plt.ylabel('Predicted Price')
plt.title('Predicted vs Actual House Prices')
plt.plot([y_test.min(), y_test.max()], [y_test.min(), y_test.max()], 'k--')
plt.show()
# Try your own prediction!
print("Enter median income to predict house price:")
user_income = float(input("Median income (for example, 3.0): "))
user_data = pd.DataFrame({
'median_income': [user_income],
'housing_median_age': [df['housing_median_age'].mean()],
'total_rooms': [df['total_rooms'].mean()],
'total_bedrooms': [df['total_bedrooms'].mean()],
'population': [df['population'].mean()],
'households': [df['households'].mean()]
})
prediction = lr.predict(user_data)
print(f"Predicted median house value: ${prediction[0]:,.2f}")
Recap and next steps#
Great work, you walked through the journey from raw data to prediction!
You learned to load, explore, clean, and model data.
You visualized and explained results.
Practice the workflow on your own with other datasets or by tweaking features.
Remember: real world projects need more cleaning, handling text columns, and trying different models.
Explore, experiment, and learn as you go. The more you practice, the better you get!
Subscribe for more beginner friendly data science videos. You are now on your data mining journey.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



