Mathew K Analytics

Lesson 34 · Python Fundamentals

Predicting California Housing Prices with Python: A Step-by-Step Regression Guide

In this lesson, you will learn the basics of Python while working on a real-world problem: predicting California house prices. You will start from the very…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb
 

Welcome to Python: California Housing Price Prediction#

In this lesson, you will learn the basics of Python while working on a real-world problem: predicting California house prices.

You will start from the very beginning and finish with a simple mini-project.

Let us get started!

# Let us start by making sure warnings will not interrupt us!
import warnings
warnings.filterwarnings('ignore')

What is Python?#

Python is a beginner-friendly programming language. It lets you solve real problems with just a few lines of code.

Today, you will use it for data analysis and prediction!

# Comments start with # and are ignored by Python.
# Use comments to explain your code.
print("Hello, California!")
Hello, California!

Variables in Python#

A variable stores information. You assign a value with = sign.

Example: price = 350000

# Let us assign some values.
location = "California"
median_price = 584000
is_expensive = True
print(location, median_price, is_expensive)
California 584000 True
# You can do basic math with Python.
a = 300000
b = 450000
total = a + b
average = total / 2
print("Total:", total)
print("Average:", average)
Total: 750000
Average: 375000.0

Lists help store multiple items#

A list is a collection of values stored in one place. You make a list with square brackets: [ ]

Example: prices = [450000, 500000, 530000]

# Creating and printing a list of values
prices = [330000, 439000, 503200, 580000]
print(prices)

# Access the first value with index 0
print(prices[0])
[330000, 439000, 503200, 580000]
330000
# Lists can hold text too
cities = ["San Diego", "Los Angeles", "San Jose"]
print(cities)

# Try printing the last city
print(cities[-1])
['San Diego', 'Los Angeles', 'San Jose']
San Jose
# Let us handle an error: What if we use a wrong index?
try:
    print(cities[5])
except IndexError:
    print("That index does not exist!")
    
That index does not exist!

The California Housing Dataset#

You will use real data about California homes. This dataset helps predict median house values in different areas.

Let us load the data and take a quick look!

# Data setup: Load California housing data
from sklearn.datasets import fetch_california_housing
import pandas as pd
cal = fetch_california_housing(as_frame=True)
df = cal.frame
print("Shape:", df.shape)
df.head()
Shape: (20640, 9)
MedInc HouseAge AveRooms AveBedrms Population AveOccup Latitude Longitude MedHouseVal
0 8.3252 41.0 6.984127 1.023810 322.0 2.555556 37.88 -122.23 4.526
1 8.3014 21.0 6.238137 0.971880 2401.0 2.109842 37.86 -122.22 3.585
2 7.2574 52.0 8.288136 1.073446 496.0 2.802260 37.85 -122.24 3.521
3 5.6431 52.0 5.817352 1.073059 558.0 2.547945 37.85 -122.25 3.413
4 3.8462 52.0 6.281853 1.081081 565.0 2.181467 37.85 -122.25 3.422
# What columns does the data have?
print(df.columns.tolist())
['MedInc', 'HouseAge', 'AveRooms', 'AveBedrms', 'Population', 'AveOccup', 'Latitude', 'Longitude', 'MedHouseVal']
# Check for missing values (empty spaces)
print(df.isnull().sum())
MedInc         0
HouseAge       0
AveRooms       0
AveBedrms      0
Population     0
AveOccup       0
Latitude       0
Longitude      0
MedHouseVal    0
dtype: int64
# Simple statistics: What is the average house value?
mean_value = df['MedHouseVal'].mean()
print("Average house value:", mean_value)
Average house value: 2.068558169089147
# Selecting data: homes worth over $3,000,000
expensive = df[df['MedHouseVal'] > 3.0]
print("Number of expensive areas:", len(expensive))
expensive.head()
Number of expensive areas: 3836
MedInc HouseAge AveRooms AveBedrms Population AveOccup Latitude Longitude MedHouseVal
0 8.3252 41.0 6.984127 1.023810 322.0 2.555556 37.88 -122.23 4.526
1 8.3014 21.0 6.238137 0.971880 2401.0 2.109842 37.86 -122.22 3.585
2 7.2574 52.0 8.288136 1.073446 496.0 2.802260 37.85 -122.24 3.521
3 5.6431 52.0 5.817352 1.073059 558.0 2.547945 37.85 -122.25 3.413
4 3.8462 52.0 6.281853 1.081081 565.0 2.181467 37.85 -122.25 3.422
# Add a new column: value divided by number of rooms
df['Value_per_room'] = df['MedHouseVal'] / df['AveRooms']
df[['MedHouseVal', 'AveRooms', 'Value_per_room']].head()
MedHouseVal AveRooms Value_per_room
0 4.526 6.984127 0.648041
1 3.585 6.238137 0.574691
2 3.521 8.288136 0.424824
3 3.413 5.817352 0.586693
4 3.422 6.281853 0.544744
# Remove the new column to clean up
del df['Value_per_room']
print('Value_per_room' in df.columns)
False
# Loop through values: first 5 house values
values = df['MedHouseVal'].head()
for val in values:
    print(val)
    
4.526
3.585
3.521
3.413
3.422
# List comprehension: double each value
nums = [1, 2, 3]
doubles = [n * 2 for n in nums]
print(doubles)
[2, 4, 6]
# Sort by median house value (lowest to highest)
sorted_df = df.sort_values('MedHouseVal')
sorted_df[['MedHouseVal']].head()
MedHouseVal
19802 0.14999
2521 0.14999
2799 0.14999
9188 0.14999
5887 0.17500
# Combine values with zip: city and value
sample_cities = ["A", "B", "C"]
sample_values = [0.5, 1.1, 2.3]
for c, v in zip(sample_cities, sample_values):
    print(c, v)
    
A 0.5
B 1.1
C 2.3
# Use input to predict: How many rooms? (simple version)
rooms = int(input("Enter average rooms in house: "))
prediction = rooms * 65000
print("Predicted value:", prediction)
Predicted value: 325000
 
# Mini-project, part 1: Prepare training and testing data
from sklearn.model_selection import train_test_split
X = df.drop('MedHouseVal', axis=1)
y = df['MedHouseVal']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
print(X_train.shape, X_test.shape)
(16512, 8) (4128, 8)
# Mini-project, part 2: Build and test a prediction model
from sklearn.linear_model import LinearRegression
model = LinearRegression()
model.fit(X_train, y_train)
preds = model.predict(X_test)
print("First 5 predictions:", preds[:5])
First 5 predictions: [0.71912284 1.76401657 2.70965883 2.83892593 2.60465725]
# How good is our model? Calculate error
from sklearn.metrics import mean_squared_error
error = mean_squared_error(y_test, preds, squared=False)
print("Root Mean Squared Error:", error)
---------------------------------------------------------------------------
TypeError                                 Traceback (most recent call last)
Cell In[22], line 3
      1 # How good is our model? Calculate error
      2 from sklearn.metrics import mean_squared_error
----> 3 error = mean_squared_error(y_test, preds, squared=False)
      4 print("Root Mean Squared Error:", error)

File c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\site-packages\sklearn\utils\_param_validation.py:196, in validate_params.<locals>.decorator.<locals>.wrapper(*args, **kwargs)
    193 func_sig = signature(func)
    195 # Map *args/**kwargs to the function signature
--> 196 params = func_sig.bind(*args, **kwargs)
    197 params.apply_defaults()
    199 # ignore self/cls and positional/keyword markers

File c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\inspect.py:3259, in Signature.bind(self, *args, **kwargs)
   3254 def bind(self, /, *args, **kwargs):
   3255     """Get a BoundArguments object, that maps the passed `args`
   3256     and `kwargs` to the function's signature.  Raises `TypeError`
   3257     if the passed arguments can not be bound.
   3258     """
-> 3259     return self._bind(args, kwargs)

File c:\Users\makmw\AppData\Local\Programs\Python\Python312\Lib\inspect.py:3248, in Signature._bind(self, args, kwargs, partial)
   3246         arguments[kwargs_param.name] = kwargs
   3247     else:
-> 3248         raise TypeError(
   3249             'got an unexpected keyword argument {arg!r}'.format(
   3250                 arg=next(iter(kwargs))))
   3252 return self._bound_arguments_cls(self, arguments)

TypeError: got an unexpected keyword argument 'squared'
# Some troubleshooting: What if our columns do not match?
try:
    wrong_X = df.drop('MedHouseVal', axis=1).drop('AveRooms', axis=1)
    model.predict(wrong_X.head())
except Exception as e:
    print("Oops! Columns in the model and new data must match.")
    
Oops! Columns in the model and new data must match.
# Tip: Use describe() to see quick stats
df.describe()
MedInc HouseAge AveRooms AveBedrms Population AveOccup Latitude Longitude MedHouseVal
count 20640.000000 20640.000000 20640.000000 20640.000000 20640.000000 20640.000000 20640.000000 20640.000000 20640.000000
mean 3.870671 28.639486 5.429000 1.096675 1425.476744 3.070655 35.631861 -119.569704 2.068558
std 1.899822 12.585558 2.474173 0.473911 1132.462122 10.386050 2.135952 2.003532 1.153956
min 0.499900 1.000000 0.846154 0.333333 3.000000 0.692308 32.540000 -124.350000 0.149990
25% 2.563400 18.000000 4.440716 1.006079 787.000000 2.429741 33.930000 -121.800000 1.196000
50% 3.534800 29.000000 5.229129 1.048780 1166.000000 2.818116 34.260000 -118.490000 1.797000
75% 4.743250 37.000000 6.052381 1.099526 1725.000000 3.282261 37.710000 -118.010000 2.647250
max 15.000100 52.000000 141.909091 34.066667 35682.000000 1243.333333 41.950000 -114.310000 5.000010
# Challenge: Predict value for a made-up house
my_data = {
    'MedInc': [6],
    'HouseAge': [30],
    'AveRooms': [8],
    'AveBedrms': [1],
    'Population': [100],
    'AveOccup': [2],
    'Latitude': [34],
    'Longitude': [-118]
}
df_input = pd.DataFrame(my_data)
print(model.predict(df_input))
[2.65440916]

Lesson Recap#

You learned Python basics and used real data to make predictions.

Great job making it to the end!

Ready for more? Practice by changing numbers and running cells again.

Thanks for following along!#

If you enjoyed this lesson, like, subscribe, and share.

Let us know in the comments which project you want next!

Happy coding!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.