Lesson 31 · Data Science Projects
Applying Linear Regression to the Boston Housing Dataset: A Step-by-Step Guide
Welcome to your beginner friendly machine learning project. Today, you will discover how to use the Boston Housing Dataset. You will learn how to set up…
- CourseData Science Projects
- Lesson31 of 33
- Video17 min
- FormatJupyter notebook · 15 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbBoston Housing Kaggle Challenge: Linear Regression Introduction#
- Welcome to your beginner friendly machine learning project.
- Today, you will discover how to use the Boston Housing Dataset.
- You will learn how to set up your data and build a basic linear regression model.
- This challenge helps teach the foundations for predicting property prices.
- No previous machine learning experience is needed.
What is the Boston Housing Dataset?#
- This dataset contains housing information for suburbs in Boston.
- It was designed for regression tasks, where you predict house prices.
- Each row is a suburb. The columns describe things like crime rate and room numbers.
- The goal is to predict the "MEDV" column, which means median house value.
- This is a famous real world benchmark for beginners to learn regression.
Why Linear Regression?#
- Linear regression is a simple and popular way to predict numbers.
- It helps you understand how features relate to target values.
- This is great for building your intuition before you try more complex methods.
# Suppress warnings for a clean output
import warnings
import numpy as np
np.random.seed(42)
warnings.filterwarnings("ignore")
# Data setup
import pandas as pd
url = 'https://raw.githubusercontent.com/selva86/datasets/master/BostonHousing.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
# Check for missing values
df.isnull().sum()
# Show column names and types
df.info()
# Quick summary statistics
df.describe()
Exploring and Understanding Features#
- Each column in the dataset is called a feature.
- Features like "crim" (crime rate) or "rm" (average number of rooms) can affect house price.
- The "medv" column is your target. That is the value you want to predict.
- Take some time to scan the feature names and use descriptions online for more details.
# Plot the target variable
import matplotlib.pyplot as plt
plt.figure(figsize=(8, 4))
plt.hist(df['medv'], bins=30, color='skyblue', edgecolor='black')
plt.title('Distribution of Median House Value (medv)')
plt.xlabel('Median Value (in $1000s)')
plt.ylabel('Frequency')
plt.show()
# Select features and target
X = df.drop('medv', axis=1)
y = df['medv']
# Split data into training and testing
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
print('Training size:', X_train.shape)
print('Testing size:', X_test.shape)
What is Linear Regression?#
- Linear regression tries to draw the best line through your data.
- It predicts a number value (for example, a price) based on the features.
- This is called a supervised learning method because you show the model what the right answer is during training.
- Linear regression is easy to use and is often the first model you try for continuous predictions.
# Build and train the linear regression model
from sklearn.linear_model import LinearRegression
lr = LinearRegression()
lr.fit(X_train, y_train)
# Make predictions on the test set
y_pred = lr.predict(X_test)
print(y_pred[:5])
# Evaluate with mean squared error and R2 score
from sklearn.metrics import mean_squared_error, r2_score
mse = mean_squared_error(y_test, y_pred)
r2 = r2_score(y_test, y_pred)
print(f'Mean Squared Error: {mse:.2f}')
print(f'R2 Score: {r2:.2f}')
# Visualize the predictions vs. the true values
plt.figure(figsize=(6, 6))
plt.scatter(y_test, y_pred, c='purple', alpha=0.6)
plt.plot([y_test.min(), y_test.max()], [y_test.min(), y_test.max()], 'k--', lw=2)
plt.xlabel('Actual MEDV')
plt.ylabel('Predicted MEDV')
plt.title('Actual vs. Predicted Median Value')
plt.show()
How to Improve Results?#
- Try using only some features or adding new ones.
- Experiment with different machine learning models.
- Tune the models parameters (these are called hyperparameters).
- Clean or transform the data for even better accuracy.
- This is what Kaggle competitions are all about: small improvements can make a big difference!
# Try your own prediction: Type values for rooms, crime rate, age, and more
# Here is a beginner interactive example
print("Let us predict a house price based on your custom inputs!")
rooms = float(input("How many rooms does the house have? "))
crime = float(input("What is the crime rate? (as a number, like 0.03) "))
age = float(input("What is the age of the house? (in years) "))
inputs = df.mean().copy()
inputs['rm'] = rooms
inputs['crim'] = crime
inputs['age'] = age
your_pred = lr.predict([inputs.drop('medv').values])[0]
print(f'Estimated median price (in $1000s): {your_pred:.2f}')
# Show coefficients to interpret what influences price
feature_importance = pd.Series(lr.coef_, index=X.columns)
feature_importance = feature_importance.sort_values(ascending=False)
print(feature_importance)
# Plot the most important features
feature_importance.head(5).plot(kind='bar', color='green')
plt.title('Top 5 Positive Feature Influences')
plt.ylabel('Coefficient Value')
plt.xlabel('Feature Name')
plt.show()
Recap and Next Steps#
- You just completed your first end to end linear regression project for house price prediction.
- You learned about regression, metrics, features, and evaluating your session.
- The skills here translate to many other prediction tasks.
- Try more datasets, experiment with features, or try tree based models for even better results.
- If this helped, please subscribe and share this video so more beginners can join.
- Happy coding and enjoy your data journey!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



