Lesson 35 · Data Science Projects
Predicting Box Office Revenue with Linear Regression: A Step-by-Step Guide
In this lesson, you will learn how to predict movie box office revenue using machine learning. Predicting revenue can help movie studios, investors, and…
- CourseData Science Projects
- Lesson35 of 33
- Video24 min
- FormatJupyter notebook · 16 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbBox Office Revenue Prediction Using Linear Regression#
- In this lesson, you will learn how to predict movie box office revenue using machine learning.
- Predicting revenue can help movie studios, investors, and data scientists make smarter decisions.
- You will explore, visualize, clean, and model movie data step by step.
- No prior experience is needed. We will explain every concept in plain language.
- Let us get started and have fun learning data mining for real world business questions!
What is Linear Regression?#
- Linear regression is a simple way to predict a number based on past patterns in data.
- In this lesson our goal is to predict box office revenue for movies.
- We use past data about movies like budget and popularity as input features.
- The model learns from examples to estimate the revenue of future movies.
- This is a supervised machine learning method. The model sees both the answers and the inputs during training.
- You will see how it works with real data and code.
# Suppress warnings for a clean notebook
import warnings; warnings.filterwarnings("ignore")
import numpy as np
np.random.seed(42)
# Data setup
import kagglehub, os, pandas as pd
path = kagglehub.dataset_download('tmdb/tmdb-movie-metadata')
files = os.listdir(path)
df = pd.read_csv(os.path.join(path, 'tmdb_5000_movies.csv'))
print(df.shape)
print(df.head(3))
# Explore basic data info
df.info()
# See column names and sample statistics
print(df.columns.tolist())
print(df.describe())
# Check for missing values in important columns
print(df[['budget', 'revenue', 'popularity']].isnull().sum())
# Remove rows where revenue or budget is zero
df = df[(df['budget'] > 0) & (df['revenue'] > 0)]
df = df.dropna(subset=['budget', 'revenue', 'popularity'])
print(df.shape)
# Check basic relationship between budget and revenue using a scatter plot
import matplotlib.pyplot as plt
plt.figure(figsize=(7,4))
plt.scatter(df['budget'], df['revenue'], alpha=0.3)
plt.xlabel('Budget')
plt.ylabel('Revenue')
plt.title('Movie Budget vs Revenue')
plt.show()
# Visualize popularity as a feature
plt.figure(figsize=(7,4))
plt.scatter(df['popularity'], df['revenue'], alpha=0.3, color='green')
plt.xlabel('Popularity Score')
plt.ylabel('Revenue')
plt.title('Popularity vs Revenue')
plt.show()
Feature Selection and Data Preparation#
- For prediction, we need to pick columns that can help explain revenue.
- We will focus on budget and popularity as our features, which are numeric and easy to use.
- Our target variable (what we want to predict) is revenue.
- We prepare and split the data before building a model.
# Create feature matrix X and target vector y
X = df[['budget', 'popularity']]
y = df['revenue']
# Split the data for training and testing
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
print('Train set size:', X_train.shape)
print('Test set size:', X_test.shape)
Building the Linear Regression Model#
- A model in data science is a way to mathematically describe patterns in data.
- Linear regression finds the line that best fits the relationship between inputs and the output.
- We will use sklearn to create and train a model.
# Fit a linear regression model
from sklearn.linear_model import LinearRegression
model = LinearRegression()
model.fit(X_train, y_train)
# Check model coefficients
print('Intercept:', model.intercept_)
print('Coefficients:', model.coef_)
# Predict on the test set
y_pred = model.predict(X_test)
print('First 5 true revenues:', y_test.values[:5])
print('First 5 predicted revenues:', y_pred[:5].astype(int))
# Evaluate performance using mean squared error and R squared
from sklearn.metrics import mean_squared_error, r2_score
mse = mean_squared_error(y_test, y_pred)
r2 = r2_score(y_test, y_pred)
print('Test set mean squared error:', mse)
print('Test set R squared:', r2)
Plotting Actual vs Predicted Revenue#
- Visualizing predictions helps us see if our model is making reasonable guesses.
- Ideally, the predicted and true values should lie close to a straight line.
# Plot predicted versus actual test revenues
plt.figure(figsize=(7,5))
plt.scatter(y_test, y_pred, alpha=0.4, color='purple')
plt.xlabel('True Revenue')
plt.ylabel('Predicted Revenue')
plt.title('Actual vs Predicted Movie Revenue')
plt.plot([y_test.min(), y_test.max()], [y_test.min(), y_test.max()], color='red', linestyle='--')
plt.show()
# Try predicting revenue for your own movie values
budget = float(input("Enter the budget (e.g. 10000000): "))
popularity = float(input("Enter the popularity score (e.g. 10): "))
prediction = model.predict([[budget, popularity]])[0]
print('Predicted box office revenue: $', int(prediction))
Practice Challenge#
- Can you add a new feature to improve prediction?
- Try using runtime or vote_average as an additional column in X.
- Retrain the model and compare the R squared score.
- Small experiments like this help you learn fast and build stronger skills.
What You Learned#
- You explored, cleaned, and visualized a real world movie dataset.
- You built and tested a simple model to predict box office revenue.
- You learned the basics of linear regression, model scores, and feature selection.
- These skills are the foundation for more advanced data mining and modeling.
- If you enjoyed this, please subscribe to our YouTube channel for more beginner friendly data science lessons!
Where to Go Next#
- Try different models like decision trees or random forest.
- Add more features or clean data differently to explore data science more deeply.
- Watch our next video to learn about classification, clustering, or recommendation systems.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



