Mathew K Analytics

Lesson 35 · Data Science Projects

Predicting Box Office Revenue with Linear Regression: A Step-by-Step Guide

In this lesson, you will learn how to predict movie box office revenue using machine learning. Predicting revenue can help movie studios, investors, and…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Box Office Revenue Prediction Using Linear Regression#

  • In this lesson, you will learn how to predict movie box office revenue using machine learning.
  • Predicting revenue can help movie studios, investors, and data scientists make smarter decisions.
  • You will explore, visualize, clean, and model movie data step by step.
  • No prior experience is needed. We will explain every concept in plain language.
  • Let us get started and have fun learning data mining for real world business questions!

What is Linear Regression?#

  • Linear regression is a simple way to predict a number based on past patterns in data.
  • In this lesson our goal is to predict box office revenue for movies.
  • We use past data about movies like budget and popularity as input features.
  • The model learns from examples to estimate the revenue of future movies.
  • This is a supervised machine learning method. The model sees both the answers and the inputs during training.
  • You will see how it works with real data and code.
# Suppress warnings for a clean notebook
import warnings; warnings.filterwarnings("ignore")
import numpy as np
np.random.seed(42)
# Data setup
import kagglehub, os, pandas as pd
path = kagglehub.dataset_download('tmdb/tmdb-movie-metadata')
files = os.listdir(path)
df = pd.read_csv(os.path.join(path, 'tmdb_5000_movies.csv'))
print(df.shape)
print(df.head(3))
(4803, 20)
      budget                                             genres  \
0  237000000  [{"id": 28, "name": "Action"}, {"id": 12, "nam...   
1  300000000  [{"id": 12, "name": "Adventure"}, {"id": 14, "...   
2  245000000  [{"id": 28, "name": "Action"}, {"id": 12, "nam...   

                                       homepage      id  \
0                   http://www.avatarmovie.com/   19995   
1  http://disney.go.com/disneypictures/pirates/     285   
2   http://www.sonypictures.com/movies/spectre/  206647   

                                            keywords original_language  \
0  [{"id": 1463, "name": "culture clash"}, {"id":...                en   
1  [{"id": 270, "name": "ocean"}, {"id": 726, "na...                en   
2  [{"id": 470, "name": "spy"}, {"id": 818, "name...                en   

                             original_title  \
0                                    Avatar   
1  Pirates of the Caribbean: At World's End   
2                                   Spectre   

                                            overview  popularity  \
0  In the 22nd century, a paraplegic Marine is di...  150.437577   
1  Captain Barbossa, long believed to be dead, ha...  139.082615   
2  A cryptic message from Bond’s past sends him o...  107.376788   

                                production_companies  \
0  [{"name": "Ingenious Film Partners", "id": 289...   
1  [{"name": "Walt Disney Pictures", "id": 2}, {"...   
2  [{"name": "Columbia Pictures", "id": 5}, {"nam...   

                                production_countries release_date     revenue  \
0  [{"iso_3166_1": "US", "name": "United States o...   2009-12-10  2787965087   
1  [{"iso_3166_1": "US", "name": "United States o...   2007-05-19   961000000   
2  [{"iso_3166_1": "GB", "name": "United Kingdom"...   2015-10-26   880674609   

   runtime                                   spoken_languages    status  \
0    162.0  [{"iso_639_1": "en", "name": "English"}, {"iso...  Released   
1    169.0           [{"iso_639_1": "en", "name": "English"}]  Released   
2    148.0  [{"iso_639_1": "fr", "name": "Fran\u00e7ais"},...  Released   

                                          tagline  \
0                     Enter the World of Pandora.   
1  At the end of the world, the adventure begins.   
2                           A Plan No One Escapes   

                                      title  vote_average  vote_count  
0                                    Avatar           7.2       11800  
1  Pirates of the Caribbean: At World's End           6.9        4500  
2                                   Spectre           6.3        4466  
# Explore basic data info
df.info()
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 4803 entries, 0 to 4802
Data columns (total 20 columns):
 #   Column                Non-Null Count  Dtype  
---  ------                --------------  -----  
 0   budget                4803 non-null   int64  
 1   genres                4803 non-null   object 
 2   homepage              1712 non-null   object 
 3   id                    4803 non-null   int64  
 4   keywords              4803 non-null   object 
 5   original_language     4803 non-null   object 
 6   original_title        4803 non-null   object 
 7   overview              4800 non-null   object 
 8   popularity            4803 non-null   float64
 9   production_companies  4803 non-null   object 
 10  production_countries  4803 non-null   object 
 11  release_date          4802 non-null   object 
 12  revenue               4803 non-null   int64  
 13  runtime               4801 non-null   float64
 14  spoken_languages      4803 non-null   object 
 15  status                4803 non-null   object 
 16  tagline               3959 non-null   object 
 17  title                 4803 non-null   object 
 18  vote_average          4803 non-null   float64
 19  vote_count            4803 non-null   int64  
dtypes: float64(3), int64(4), object(13)
memory usage: 750.6+ KB
# See column names and sample statistics
print(df.columns.tolist())
print(df.describe())
['budget', 'genres', 'homepage', 'id', 'keywords', 'original_language', 'original_title', 'overview', 'popularity', 'production_companies', 'production_countries', 'release_date', 'revenue', 'runtime', 'spoken_languages', 'status', 'tagline', 'title', 'vote_average', 'vote_count']
             budget             id   popularity       revenue      runtime  \
count  4.803000e+03    4803.000000  4803.000000  4.803000e+03  4801.000000   
mean   2.904504e+07   57165.484281    21.492301  8.226064e+07   106.875859   
std    4.072239e+07   88694.614033    31.816650  1.628571e+08    22.611935   
min    0.000000e+00       5.000000     0.000000  0.000000e+00     0.000000   
25%    7.900000e+05    9014.500000     4.668070  0.000000e+00    94.000000   
50%    1.500000e+07   14629.000000    12.921594  1.917000e+07   103.000000   
75%    4.000000e+07   58610.500000    28.313505  9.291719e+07   118.000000   
max    3.800000e+08  459488.000000   875.581305  2.787965e+09   338.000000   

       vote_average    vote_count  
count   4803.000000   4803.000000  
mean       6.092172    690.217989  
std        1.194612   1234.585891  
min        0.000000      0.000000  
25%        5.600000     54.000000  
50%        6.200000    235.000000  
75%        6.800000    737.000000  
max       10.000000  13752.000000  
# Check for missing values in important columns
print(df[['budget', 'revenue', 'popularity']].isnull().sum())
budget        0
revenue       0
popularity    0
dtype: int64
# Remove rows where revenue or budget is zero
df = df[(df['budget'] > 0) & (df['revenue'] > 0)]
df = df.dropna(subset=['budget', 'revenue', 'popularity'])
print(df.shape)
(3229, 20)
# Check basic relationship between budget and revenue using a scatter plot
import matplotlib.pyplot as plt
plt.figure(figsize=(7,4))
plt.scatter(df['budget'], df['revenue'], alpha=0.3)
plt.xlabel('Budget')
plt.ylabel('Revenue')
plt.title('Movie Budget vs Revenue')
plt.show()
No description has been provided for this image
# Visualize popularity as a feature
plt.figure(figsize=(7,4))
plt.scatter(df['popularity'], df['revenue'], alpha=0.3, color='green')
plt.xlabel('Popularity Score')
plt.ylabel('Revenue')
plt.title('Popularity vs Revenue')
plt.show()
No description has been provided for this image

Feature Selection and Data Preparation#

  • For prediction, we need to pick columns that can help explain revenue.
  • We will focus on budget and popularity as our features, which are numeric and easy to use.
  • Our target variable (what we want to predict) is revenue.
  • We prepare and split the data before building a model.
# Create feature matrix X and target vector y
X = df[['budget', 'popularity']]
y = df['revenue']
# Split the data for training and testing
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
print('Train set size:', X_train.shape)
print('Test set size:', X_test.shape)
Train set size: (2583, 2)
Test set size: (646, 2)

Building the Linear Regression Model#

  • A model in data science is a way to mathematically describe patterns in data.
  • Linear regression finds the line that best fits the relationship between inputs and the output.
  • We will use sklearn to create and train a model.
# Fit a linear regression model
from sklearn.linear_model import LinearRegression
model = LinearRegression()
model.fit(X_train, y_train)
LinearRegression()
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
# Check model coefficients
print('Intercept:', model.intercept_)
print('Coefficients:', model.coef_)
Intercept: -21720322.659342214
Coefficients: [2.20111951e+00 1.80922061e+06]
# Predict on the test set
y_pred = model.predict(X_test)
print('First 5 true revenues:', y_test.values[:5])
print('First 5 predicted revenues:', y_pred[:5].astype(int))
First 5 true revenues: [169852759  96408652 140073390   7000000  68296293]
First 5 predicted revenues: [204055212  70563378 221892694   1746775  61202841]
# Evaluate performance using mean squared error and R squared
from sklearn.metrics import mean_squared_error, r2_score
mse = mean_squared_error(y_test, y_pred)
r2 = r2_score(y_test, y_pred)
print('Test set mean squared error:', mse)
print('Test set R squared:', r2)
Test set mean squared error: 2.2501400109522508e+16
Test set R squared: 0.5548262651995053

Plotting Actual vs Predicted Revenue#

  • Visualizing predictions helps us see if our model is making reasonable guesses.
  • Ideally, the predicted and true values should lie close to a straight line.
# Plot predicted versus actual test revenues
plt.figure(figsize=(7,5))
plt.scatter(y_test, y_pred, alpha=0.4, color='purple')
plt.xlabel('True Revenue')
plt.ylabel('Predicted Revenue')
plt.title('Actual vs Predicted Movie Revenue')
plt.plot([y_test.min(), y_test.max()], [y_test.min(), y_test.max()], color='red', linestyle='--')
plt.show()
No description has been provided for this image
# Try predicting revenue for your own movie values
budget = float(input("Enter the budget (e.g. 10000000): "))
popularity = float(input("Enter the popularity score (e.g. 10): "))
prediction = model.predict([[budget, popularity]])[0]
print('Predicted box office revenue: $', int(prediction))
Predicted box office revenue: $ 115473961

Practice Challenge#

  • Can you add a new feature to improve prediction?
  • Try using runtime or vote_average as an additional column in X.
  • Retrain the model and compare the R squared score.
  • Small experiments like this help you learn fast and build stronger skills.

What You Learned#

  • You explored, cleaned, and visualized a real world movie dataset.
  • You built and tested a simple model to predict box office revenue.
  • You learned the basics of linear regression, model scores, and feature selection.
  • These skills are the foundation for more advanced data mining and modeling.
  • If you enjoyed this, please subscribe to our YouTube channel for more beginner friendly data science lessons!

Where to Go Next#

  • Try different models like decision trees or random forest.
  • Add more features or clean data differently to explore data science more deeply.
  • Watch our next video to learn about classification, clustering, or recommendation systems.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.