Mathew K Analytics

Lesson 30 · Data Science Projects

Building a House Price Prediction Model with Python: Data Science and Machine Learning

In this lesson, you will learn how to predict house prices using real world data. You will see the full process: loading data, inspecting it, doing…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

House Price Prediction using Machine Learning#

  • In this lesson, you will learn how to predict house prices using real world data.

  • You will see the full process: loading data, inspecting it, doing analysis, cleaning, building a model, and making predictions.

  • House price prediction is a classic example in data mining and real estate analytics.

  • You will also learn to visualize, explain, and improve model results.

  • This is for absolute beginners. No previous coding or machine learning experience is needed.

  • Let us get started and discover the power of Python in data science!

Lesson Goals#

  • Understand the process of predicting house prices using data.

  • Explore and prepare real house price data with pandas.

  • Build and train a simple machine learning model.

  • Evaluate model accuracy and make predictions.

  • Learn to visualize data and results.

  • By the end, you will be able to apply this workflow to other datasets.

  • Remember to subscribe if you enjoy these beginner lessons!

# Suppress warnings to make outputs clean
import warnings; warnings.filterwarnings("ignore")
import numpy as np
np.random.seed(42)

Data setup#

  • We are using the California Housing Dataset for this lesson.
  • This data is real and describes housing in California districts.
  • It includes features like number of rooms, population, and median value.
  • We will load the data with pandas and preview its shape and contents.
# Load the California housing dataset
import pandas as pd
url = 'https://raw.githubusercontent.com/ageron/handson-ml/master/datasets/housing/housing.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
(20640, 10)
   longitude  latitude  housing_median_age  total_rooms  total_bedrooms  \
0    -122.23     37.88                41.0        880.0           129.0   
1    -122.22     37.86                21.0       7099.0          1106.0   
2    -122.24     37.85                52.0       1467.0           190.0   

   population  households  median_income  median_house_value ocean_proximity  
0       322.0       126.0         8.3252            452600.0        NEAR BAY  
1      2401.0      1138.0         8.3014            358500.0        NEAR BAY  
2       496.0       177.0         7.2574            352100.0        NEAR BAY  
# Let us check the column names and types
print(df.dtypes)
longitude             float64
latitude              float64
housing_median_age    float64
total_rooms           float64
total_bedrooms        float64
population            float64
households            float64
median_income         float64
median_house_value    float64
ocean_proximity        object
dtype: object
# Summary statistics help us see data shape
print(df.describe())
          longitude      latitude  housing_median_age   total_rooms  \
count  20640.000000  20640.000000        20640.000000  20640.000000   
mean    -119.569704     35.631861           28.639486   2635.763081   
std        2.003532      2.135952           12.585558   2181.615252   
min     -124.350000     32.540000            1.000000      2.000000   
25%     -121.800000     33.930000           18.000000   1447.750000   
50%     -118.490000     34.260000           29.000000   2127.000000   
75%     -118.010000     37.710000           37.000000   3148.000000   
max     -114.310000     41.950000           52.000000  39320.000000   

       total_bedrooms    population    households  median_income  \
count    20433.000000  20640.000000  20640.000000   20640.000000   
mean       537.870553   1425.476744    499.539680       3.870671   
std        421.385070   1132.462122    382.329753       1.899822   
min          1.000000      3.000000      1.000000       0.499900   
25%        296.000000    787.000000    280.000000       2.563400   
50%        435.000000   1166.000000    409.000000       3.534800   
75%        647.000000   1725.000000    605.000000       4.743250   
max       6445.000000  35682.000000   6082.000000      15.000100   

       median_house_value  
count        20640.000000  
mean        206855.816909  
std         115395.615874  
min          14999.000000  
25%         119600.000000  
50%         179700.000000  
75%         264725.000000  
max         500001.000000  

Checking for missing data#

  • Real world data is not always perfect.
  • Let us check whether we have any missing values in our dataset.
  • Missing values can cause problems when building models, so it is best to catch them early.
# Find missing values by column
print(df.isnull().sum())
longitude               0
latitude                0
housing_median_age      0
total_rooms             0
total_bedrooms        207
population              0
households              0
median_income           0
median_house_value      0
ocean_proximity         0
dtype: int64
# Remove rows with missing values
df = df.dropna()
print(df.shape)
(20433, 10)

Visualizing the target variable#

  • The target variable is what we want to predict. In this dataset, it is 'median_house_value'.
  • Let us look at its distribution to understand its range and spread.
  • Visualizing the target helps spot any extreme or unusual values.
# Plot the distribution of median house values
import matplotlib.pyplot as plt
plt.figure(figsize=(8,4))
plt.hist(df['median_house_value'], bins=30, color='skyblue')
plt.title('Distribution of Median House Value')
plt.xlabel('Median House Value')
plt.ylabel('Number of Districts')
plt.show()
No description has been provided for this image
# Quick look at correlation with house value
corr = df.corr()
print(corr['median_house_value'].sort_values(ascending=False))
---------------------------------------------------------------------------
ValueError                                Traceback (most recent call last)
Cell In[8], line 2
      1 # Quick look at correlation with house value
----> 2 corr = df.corr()
      3 print(corr['median_house_value'].sort_values(ascending=False))

File c:\Python\Lib\site-packages\pandas\core\frame.py:11049, in DataFrame.corr(self, method, min_periods, numeric_only)
  11047 cols = data.columns
  11048 idx = cols.copy()
> 11049 mat = data.to_numpy(dtype=float, na_value=np.nan, copy=False)
  11051 if method == "pearson":
  11052     correl = libalgos.nancorr(mat, minp=min_periods)

File c:\Python\Lib\site-packages\pandas\core\frame.py:1993, in DataFrame.to_numpy(self, dtype, copy, na_value)
   1991 if dtype is not None:
   1992     dtype = np.dtype(dtype)
-> 1993 result = self._mgr.as_array(dtype=dtype, copy=copy, na_value=na_value)
   1994 if result.dtype is not dtype:
   1995     result = np.asarray(result, dtype=dtype)

File c:\Python\Lib\site-packages\pandas\core\internals\managers.py:1694, in BlockManager.as_array(self, dtype, copy, na_value)
   1692         arr.flags.writeable = False
   1693 else:
-> 1694     arr = self._interleave(dtype=dtype, na_value=na_value)
   1695     # The underlying data was copied within _interleave, so no need
   1696     # to further copy if copy=True or setting na_value
   1698 if na_value is lib.no_default:

File c:\Python\Lib\site-packages\pandas\core\internals\managers.py:1753, in BlockManager._interleave(self, dtype, na_value)
   1751     else:
   1752         arr = blk.get_values(dtype)
-> 1753     result[rl.indexer] = arr
   1754     itemmask[rl.indexer] = 1
   1756 if not itemmask.all():

ValueError: could not convert string to float: 'NEAR BAY'
# Visualize longitude and latitude
plt.figure(figsize=(8,6))
plt.scatter(df['longitude'], df['latitude'], alpha=0.2,
            c=df['median_house_value'], cmap='viridis')
plt.colorbar(label='Median House Value')
plt.xlabel('Longitude')
plt.ylabel('Latitude')
plt.title('House Prices by Location')
plt.show()
No description has been provided for this image

Feature selection#

  • To build a prediction model, we need to choose which features to use.

  • For this lesson, we will use simple numeric columns.

  • Picking the right features can make models perform better and be easier to understand.

# Select numeric features
features = [
    'median_income',
    'housing_median_age',
    'total_rooms',
    'total_bedrooms',
    'population',
    'households'
]
X = df[features]
y = df['median_house_value']
print(X.shape, y.shape)
(20433, 6) (20433,)
# Split into training and test sets
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42)
print(X_train.shape, X_test.shape)
(16346, 6) (4087, 6)

Building a linear regression model#

  • A linear regression model is a simple and popular choice for predicting numbers.
  • It learns to draw a straight line that best fits the training data.
  • Let us build and train the model using scikit learn, a friendly Python package.
# Train a linear regression model
from sklearn.linear_model import LinearRegression
lr = LinearRegression()
lr.fit(X_train, y_train)
LinearRegression()
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
# Predict prices for test data
y_pred = lr.predict(X_test)
print(y_pred[:5])
[165480.22752968 161093.72645381 196854.06878554 154256.54168805
 199116.5340472 ]
# Measure model accuracy
from sklearn.metrics import mean_squared_error, r2_score
mse = mean_squared_error(y_test, y_pred)
rmse = mse ** 0.5
r2 = r2_score(y_test, y_pred)
print(f"Root Mean Squared Error: {rmse:.2f}")
print(f"R2 Score: {r2:.3f}")
Root Mean Squared Error: 76587.33
R2 Score: 0.571
# Plot predicted vs actual house prices
plt.figure(figsize=(6,6))
plt.scatter(y_test, y_pred, alpha=0.2, color='purple')
plt.xlabel('Actual Price')
plt.ylabel('Predicted Price')
plt.title('Predicted vs Actual House Prices')
plt.plot([y_test.min(), y_test.max()], [y_test.min(), y_test.max()], 'k--')
plt.show()
No description has been provided for this image
# Try your own prediction!
print("Enter median income to predict house price:")
user_income = float(input("Median income (for example, 3.0): "))
user_data = pd.DataFrame({
    'median_income': [user_income],
    'housing_median_age': [df['housing_median_age'].mean()],
    'total_rooms': [df['total_rooms'].mean()],
    'total_bedrooms': [df['total_bedrooms'].mean()],
    'population': [df['population'].mean()],
    'households': [df['households'].mean()]
})
prediction = lr.predict(user_data)
print(f"Predicted median house value: ${prediction[0]:,.2f}")
Enter median income to predict house price:
Predicted median house value: $308,296.41

Recap and next steps#

  • Great work, you walked through the journey from raw data to prediction!

  • You learned to load, explore, clean, and model data.

  • You visualized and explained results.

  • Practice the workflow on your own with other datasets or by tweaking features.

  • Remember: real world projects need more cleaning, handling text columns, and trying different models.

  • Explore, experiment, and learn as you go. The more you practice, the better you get!

  • Subscribe for more beginner friendly data science videos. You are now on your data mining journey.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.