Mathew K Analytics

Lesson 31 · Data Science Projects

Applying Linear Regression to the Boston Housing Dataset: A Step-by-Step Guide

Welcome to your beginner friendly machine learning project. Today, you will discover how to use the Boston Housing Dataset. You will learn how to set up…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Boston Housing Kaggle Challenge: Linear Regression Introduction#

  • Welcome to your beginner friendly machine learning project.
  • Today, you will discover how to use the Boston Housing Dataset.
  • You will learn how to set up your data and build a basic linear regression model.
  • This challenge helps teach the foundations for predicting property prices.
  • No previous machine learning experience is needed.

What is the Boston Housing Dataset?#

  • This dataset contains housing information for suburbs in Boston.
  • It was designed for regression tasks, where you predict house prices.
  • Each row is a suburb. The columns describe things like crime rate and room numbers.
  • The goal is to predict the "MEDV" column, which means median house value.
  • This is a famous real world benchmark for beginners to learn regression.

Why Linear Regression?#

  • Linear regression is a simple and popular way to predict numbers.
  • It helps you understand how features relate to target values.
  • This is great for building your intuition before you try more complex methods.
# Suppress warnings for a clean output
import warnings
import numpy as np
np.random.seed(42)
warnings.filterwarnings("ignore")
# Data setup
import pandas as pd
url = 'https://raw.githubusercontent.com/selva86/datasets/master/BostonHousing.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
(506, 14)
      crim    zn  indus  chas    nox     rm   age     dis  rad  tax  ptratio  \
0  0.00632  18.0   2.31     0  0.538  6.575  65.2  4.0900    1  296     15.3   
1  0.02731   0.0   7.07     0  0.469  6.421  78.9  4.9671    2  242     17.8   
2  0.02729   0.0   7.07     0  0.469  7.185  61.1  4.9671    2  242     17.8   

        b  lstat  medv  
0  396.90   4.98  24.0  
1  396.90   9.14  21.6  
2  392.83   4.03  34.7  
# Check for missing values
df.isnull().sum()
crim       0
zn         0
indus      0
chas       0
nox        0
rm         0
age        0
dis        0
rad        0
tax        0
ptratio    0
b          0
lstat      0
medv       0
dtype: int64
# Show column names and types
df.info()
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 506 entries, 0 to 505
Data columns (total 14 columns):
 #   Column   Non-Null Count  Dtype  
---  ------   --------------  -----  
 0   crim     506 non-null    float64
 1   zn       506 non-null    float64
 2   indus    506 non-null    float64
 3   chas     506 non-null    int64  
 4   nox      506 non-null    float64
 5   rm       506 non-null    float64
 6   age      506 non-null    float64
 7   dis      506 non-null    float64
 8   rad      506 non-null    int64  
 9   tax      506 non-null    int64  
 10  ptratio  506 non-null    float64
 11  b        506 non-null    float64
 12  lstat    506 non-null    float64
 13  medv     506 non-null    float64
dtypes: float64(11), int64(3)
memory usage: 55.5 KB
# Quick summary statistics
df.describe()
crim zn indus chas nox rm age dis rad tax ptratio b lstat medv
count 506.000000 506.000000 506.000000 506.000000 506.000000 506.000000 506.000000 506.000000 506.000000 506.000000 506.000000 506.000000 506.000000 506.000000
mean 3.613524 11.363636 11.136779 0.069170 0.554695 6.284634 68.574901 3.795043 9.549407 408.237154 18.455534 356.674032 12.653063 22.532806
std 8.601545 23.322453 6.860353 0.253994 0.115878 0.702617 28.148861 2.105710 8.707259 168.537116 2.164946 91.294864 7.141062 9.197104
min 0.006320 0.000000 0.460000 0.000000 0.385000 3.561000 2.900000 1.129600 1.000000 187.000000 12.600000 0.320000 1.730000 5.000000
25% 0.082045 0.000000 5.190000 0.000000 0.449000 5.885500 45.025000 2.100175 4.000000 279.000000 17.400000 375.377500 6.950000 17.025000
50% 0.256510 0.000000 9.690000 0.000000 0.538000 6.208500 77.500000 3.207450 5.000000 330.000000 19.050000 391.440000 11.360000 21.200000
75% 3.677083 12.500000 18.100000 0.000000 0.624000 6.623500 94.075000 5.188425 24.000000 666.000000 20.200000 396.225000 16.955000 25.000000
max 88.976200 100.000000 27.740000 1.000000 0.871000 8.780000 100.000000 12.126500 24.000000 711.000000 22.000000 396.900000 37.970000 50.000000

Exploring and Understanding Features#

  • Each column in the dataset is called a feature.
  • Features like "crim" (crime rate) or "rm" (average number of rooms) can affect house price.
  • The "medv" column is your target. That is the value you want to predict.
  • Take some time to scan the feature names and use descriptions online for more details.
# Plot the target variable
import matplotlib.pyplot as plt
plt.figure(figsize=(8, 4))
plt.hist(df['medv'], bins=30, color='skyblue', edgecolor='black')
plt.title('Distribution of Median House Value (medv)')
plt.xlabel('Median Value (in $1000s)')
plt.ylabel('Frequency')
plt.show()
No description has been provided for this image
# Select features and target
X = df.drop('medv', axis=1)
y = df['medv']
# Split data into training and testing
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
print('Training size:', X_train.shape)
print('Testing size:', X_test.shape)
Training size: (404, 13)
Testing size: (102, 13)

What is Linear Regression?#

  • Linear regression tries to draw the best line through your data.
  • It predicts a number value (for example, a price) based on the features.
  • This is called a supervised learning method because you show the model what the right answer is during training.
  • Linear regression is easy to use and is often the first model you try for continuous predictions.
# Build and train the linear regression model
from sklearn.linear_model import LinearRegression
lr = LinearRegression()
lr.fit(X_train, y_train)
LinearRegression()
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
# Make predictions on the test set
y_pred = lr.predict(X_test)
print(y_pred[:5])
[28.99672362 36.02556534 14.81694405 25.03197915 18.76987992]
# Evaluate with mean squared error and R2 score
from sklearn.metrics import mean_squared_error, r2_score
mse = mean_squared_error(y_test, y_pred)
r2 = r2_score(y_test, y_pred)
print(f'Mean Squared Error: {mse:.2f}')
print(f'R2 Score: {r2:.2f}')
Mean Squared Error: 24.29
R2 Score: 0.67
# Visualize the predictions vs. the true values
plt.figure(figsize=(6, 6))
plt.scatter(y_test, y_pred, c='purple', alpha=0.6)
plt.plot([y_test.min(), y_test.max()], [y_test.min(), y_test.max()], 'k--', lw=2)
plt.xlabel('Actual MEDV')
plt.ylabel('Predicted MEDV')
plt.title('Actual vs. Predicted Median Value')
plt.show()
No description has been provided for this image

How to Improve Results?#

  • Try using only some features or adding new ones.
  • Experiment with different machine learning models.
  • Tune the models parameters (these are called hyperparameters).
  • Clean or transform the data for even better accuracy.
  • This is what Kaggle competitions are all about: small improvements can make a big difference!
# Try your own prediction: Type values for rooms, crime rate, age, and more
# Here is a beginner interactive example
print("Let us predict a house price based on your custom inputs!")
rooms = float(input("How many rooms does the house have? "))
crime = float(input("What is the crime rate? (as a number, like 0.03) "))
age = float(input("What is the age of the house? (in years) "))
inputs = df.mean().copy()
inputs['rm'] = rooms
inputs['crim'] = crime
inputs['age'] = age
your_pred = lr.predict([inputs.drop('medv').values])[0]
print(f'Estimated median price (in $1000s): {your_pred:.2f}')
Let us predict a house price based on your custom inputs!
Estimated median price (in $1000s): 21.87
# Show coefficients to interpret what influences price
feature_importance = pd.Series(lr.coef_, index=X.columns)
feature_importance = feature_importance.sort_values(ascending=False)
print(feature_importance)
rm          4.438835
chas        2.784438
rad         0.262430
indus       0.040381
zn          0.030110
b           0.012351
age        -0.006296
tax        -0.010647
crim       -0.113056
lstat      -0.508571
ptratio    -0.915456
dis        -1.447865
nox       -17.202633
dtype: float64
# Plot the most important features
feature_importance.head(5).plot(kind='bar', color='green')
plt.title('Top 5 Positive Feature Influences')
plt.ylabel('Coefficient Value')
plt.xlabel('Feature Name')
plt.show()
No description has been provided for this image

Recap and Next Steps#

  • You just completed your first end to end linear regression project for house price prediction.
  • You learned about regression, metrics, features, and evaluating your session.
  • The skills here translate to many other prediction tasks.
  • Try more datasets, experiment with features, or try tree based models for even better results.
  • If this helped, please subscribe and share this video so more beginners can join.
  • Happy coding and enjoy your data journey!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.