Lesson 34 · Python Fundamentals
Predicting California Housing Prices with Python: A Step-by-Step Regression Guide
In this lesson, you will learn the basics of Python while working on a real-world problem: predicting California house prices. You will start from the very…
- CoursePython Fundamentals
- Lesson34 of 22
- Video16 min
- FormatJupyter notebook · 27 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynb
Welcome to Python: California Housing Price Prediction#
In this lesson, you will learn the basics of Python while working on a real-world problem: predicting California house prices.
You will start from the very beginning and finish with a simple mini-project.
Let us get started!
# Let us start by making sure warnings will not interrupt us!
import warnings
warnings.filterwarnings('ignore')
What is Python?#
Python is a beginner-friendly programming language. It lets you solve real problems with just a few lines of code.
Today, you will use it for data analysis and prediction!
# Comments start with # and are ignored by Python.
# Use comments to explain your code.
print("Hello, California!")
Variables in Python#
A variable stores information. You assign a value with = sign.
Example: price = 350000
# Let us assign some values.
location = "California"
median_price = 584000
is_expensive = True
print(location, median_price, is_expensive)
# You can do basic math with Python.
a = 300000
b = 450000
total = a + b
average = total / 2
print("Total:", total)
print("Average:", average)
Lists help store multiple items#
A list is a collection of values stored in one place. You make a list with square brackets: [ ]
Example: prices = [450000, 500000, 530000]
# Creating and printing a list of values
prices = [330000, 439000, 503200, 580000]
print(prices)
# Access the first value with index 0
print(prices[0])
# Lists can hold text too
cities = ["San Diego", "Los Angeles", "San Jose"]
print(cities)
# Try printing the last city
print(cities[-1])
# Let us handle an error: What if we use a wrong index?
try:
print(cities[5])
except IndexError:
print("That index does not exist!")
The California Housing Dataset#
You will use real data about California homes. This dataset helps predict median house values in different areas.
Let us load the data and take a quick look!
# Data setup: Load California housing data
from sklearn.datasets import fetch_california_housing
import pandas as pd
cal = fetch_california_housing(as_frame=True)
df = cal.frame
print("Shape:", df.shape)
df.head()
# What columns does the data have?
print(df.columns.tolist())
# Check for missing values (empty spaces)
print(df.isnull().sum())
# Simple statistics: What is the average house value?
mean_value = df['MedHouseVal'].mean()
print("Average house value:", mean_value)
# Selecting data: homes worth over $3,000,000
expensive = df[df['MedHouseVal'] > 3.0]
print("Number of expensive areas:", len(expensive))
expensive.head()
# Add a new column: value divided by number of rooms
df['Value_per_room'] = df['MedHouseVal'] / df['AveRooms']
df[['MedHouseVal', 'AveRooms', 'Value_per_room']].head()
# Remove the new column to clean up
del df['Value_per_room']
print('Value_per_room' in df.columns)
# Loop through values: first 5 house values
values = df['MedHouseVal'].head()
for val in values:
print(val)
# List comprehension: double each value
nums = [1, 2, 3]
doubles = [n * 2 for n in nums]
print(doubles)
# Sort by median house value (lowest to highest)
sorted_df = df.sort_values('MedHouseVal')
sorted_df[['MedHouseVal']].head()
# Combine values with zip: city and value
sample_cities = ["A", "B", "C"]
sample_values = [0.5, 1.1, 2.3]
for c, v in zip(sample_cities, sample_values):
print(c, v)
# Use input to predict: How many rooms? (simple version)
rooms = int(input("Enter average rooms in house: "))
prediction = rooms * 65000
print("Predicted value:", prediction)
# Mini-project, part 1: Prepare training and testing data
from sklearn.model_selection import train_test_split
X = df.drop('MedHouseVal', axis=1)
y = df['MedHouseVal']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
print(X_train.shape, X_test.shape)
# Mini-project, part 2: Build and test a prediction model
from sklearn.linear_model import LinearRegression
model = LinearRegression()
model.fit(X_train, y_train)
preds = model.predict(X_test)
print("First 5 predictions:", preds[:5])
# How good is our model? Calculate error
from sklearn.metrics import mean_squared_error
error = mean_squared_error(y_test, preds, squared=False)
print("Root Mean Squared Error:", error)
# Some troubleshooting: What if our columns do not match?
try:
wrong_X = df.drop('MedHouseVal', axis=1).drop('AveRooms', axis=1)
model.predict(wrong_X.head())
except Exception as e:
print("Oops! Columns in the model and new data must match.")
# Tip: Use describe() to see quick stats
df.describe()
# Challenge: Predict value for a made-up house
my_data = {
'MedInc': [6],
'HouseAge': [30],
'AveRooms': [8],
'AveBedrms': [1],
'Population': [100],
'AveOccup': [2],
'Latitude': [34],
'Longitude': [-118]
}
df_input = pd.DataFrame(my_data)
print(model.predict(df_input))
Lesson Recap#
You learned Python basics and used real data to make predictions.
Great job making it to the end!
Ready for more? Practice by changing numbers and running cells again.
Thanks for following along!#
If you enjoyed this lesson, like, subscribe, and share.
Let us know in the comments which project you want next!
Happy coding!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



