Mathew K Analytics

Lesson 41 · Data Science Projects

Building a Recommendation System in Python: A Step-by-Step Guide for Beginners

Learn how to build simple recommendation models that suggest items for users. Recommendation systems are used by companies like Netflix, Amazon, and…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Recommendation System in Python#

  • Learn how to build simple recommendation models that suggest items for users.
  • Recommendation systems are used by companies like Netflix, Amazon, and YouTube.
  • You will discover the basic logic of machine learning for making personalized suggestions.
  • This topic matters because smart recommendations help users find what they need quickly.
  • Let us get started with data and a real example!

Real World Dataset: MovieLens#

  • We use the MovieLens Ratings Dataset, a classic source for building and testing recommender systems.
  • It contains user ids, movie ids, and ratings from real people.
  • This is similar to how Netflix tracks what shows you like.
# Suppress warnings for a clean notebook
import warnings; warnings.filterwarnings("ignore")
import numpy as np
np.random.seed(42)
 
# Data setup
import pandas as pd, zipfile, requests, io
url = 'https://files.grouplens.org/datasets/movielens/ml-latest-small.zip'
r = requests.get(url)
z = zipfile.ZipFile(io.BytesIO(r.content))
df = pd.read_csv(z.open('ml-latest-small/ratings.csv'))
print(df.shape)
print(df.head(3))
(100836, 4)
   userId  movieId  rating  timestamp
0       1        1     4.0  964982703
1       1        3     4.0  964981247
2       1        6     4.0  964982224
# Check basic info about our data
df.info()
 
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 100836 entries, 0 to 100835
Data columns (total 4 columns):
 #   Column     Non-Null Count   Dtype  
---  ------     --------------   -----  
 0   userId     100836 non-null  int64  
 1   movieId    100836 non-null  int64  
 2   rating     100836 non-null  float64
 3   timestamp  100836 non-null  int64  
dtypes: float64(1), int64(3)
memory usage: 3.1 MB
# Let us look at the unique users and movies
num_users = df['userId'].nunique()
num_movies = df['movieId'].nunique()
print(f"There are {num_users} unique users and {num_movies} unique movies.")
There are 610 unique users and 9724 unique movies.

What Is a Recommendation System?#

  • A recommendation system is a tool that suggests items (like movies or products) to users based on their interests.
  • There are two main types: content-based and collaborative filtering.
  • We will start with collaborative filtering because it uses patterns of ratings from many users.
  • This is like finding users who are similar to you and learning what they liked.
# Let us find the average rating given by each user
user_means = df.groupby('userId')['rating'].mean()
print(user_means.head())
userId
1    4.366379
2    3.948276
3    2.435897
4    3.555556
5    3.636364
Name: rating, dtype: float64
# Let us explore ratings for a specific movie
movie_stats = df.groupby('movieId')['rating'].agg(['count', 'mean']).sort_values(by='count', ascending=False)
print(movie_stats.head(5))
         count      mean
movieId                 
356        329  4.164134
318        317  4.429022
296        307  4.197068
593        279  4.161290
2571       278  4.192446
# Create a user-movie rating matrix
rating_matrix = df.pivot_table(index='userId', columns='movieId', values='rating')
print(rating_matrix.shape)
print(rating_matrix.head(3))
(610, 9724)
movieId  1       2       3       4       5       6       7       8       \
userId                                                                    
1           4.0     NaN     4.0     NaN     NaN     4.0     NaN     NaN   
2           NaN     NaN     NaN     NaN     NaN     NaN     NaN     NaN   
3           NaN     NaN     NaN     NaN     NaN     NaN     NaN     NaN   

movieId  9       10      ...  193565  193567  193571  193573  193579  193581  \
userId                   ...                                                   
1           NaN     NaN  ...     NaN     NaN     NaN     NaN     NaN     NaN   
2           NaN     NaN  ...     NaN     NaN     NaN     NaN     NaN     NaN   
3           NaN     NaN  ...     NaN     NaN     NaN     NaN     NaN     NaN   

movieId  193583  193585  193587  193609  
userId                                   
1           NaN     NaN     NaN     NaN  
2           NaN     NaN     NaN     NaN  
3           NaN     NaN     NaN     NaN  

[3 rows x 9724 columns]
# Visualize the distribution of ratings
import matplotlib.pyplot as plt
import seaborn as sns
sns.histplot(df['rating'], bins=10, kde=True)
plt.xlabel('Rating Value')
plt.title('Distribution of Movie Ratings')
plt.show()
No description has been provided for this image
# Let us try user similarity: compute cosine similarity
from sklearn.metrics.pairwise import cosine_similarity
user_similarity = cosine_similarity(rating_matrix.fillna(0))
print(user_similarity.shape)
print(user_similarity[:3, :3])
(610, 610)
[[1.         0.02728287 0.05972026]
 [0.02728287 1.         0.        ]
 [0.05972026 0.         1.        ]]
# Pick a random user to recommend movies for
user_id = df['userId'].sample().values[0]
print(f"Selected user id: {user_id}")
Selected user id: 432
# Find the most similar users to our target user
target_idx = rating_matrix.index.get_loc(user_id)
sim_scores = user_similarity[target_idx]
similar_users = rating_matrix.index[sim_scores.argsort()[::-1][1:6]]
print(f"Most similar users: {list(similar_users)}")
Most similar users: [434, 239, 247, 573, 254]
# Gather movies liked by similar users but not yet rated by our user
user_movies = set(df[df['userId'] == user_id]['movieId'])
sim_movies = df[df['userId'].isin(similar_users) & ~df['movieId'].isin(user_movies)]
top_movies = sim_movies.groupby('movieId')['rating'].mean().sort_values(ascending=False).head(5)
print("Recommended movies:", list(top_movies.index))
Recommended movies: [4816, 96079, 92535, 3213, 88125]
# Show a mini project: build for a different user
another_user = int(input("Enter a user id to get movie recommendations (pick a number from your earlier list): "))
target_idx = rating_matrix.index.get_loc(another_user)
sim_scores = user_similarity[target_idx]
similar_users = rating_matrix.index[sim_scores.argsort()[::-1][1:6]]
user_movies = set(df[df['userId'] == another_user]['movieId'])
sim_movies = df[df['userId'].isin(similar_users) & ~df['movieId'].isin(user_movies)]
top_movies = sim_movies.groupby('movieId')['rating'].mean().sort_values(ascending=False).head(5)
print("Recommendations for your user:", list(top_movies.index))
Recommendations for your user: [47099, 2712, 92259, 48780, 84152]

Behind the Scenes: How Does Collaborative Filtering Work?#

  • These recommendations are based on what similar people preferred.
  • In the real world, extra steps such as normalization, more advanced models, or using movie titles are common.
  • But you have learned the core idea that powers many streaming websites.
# Optional: Map movie ids back to titles for better user experience
movies_url = 'https://files.grouplens.org/datasets/movielens/ml-latest-small.zip'
r = requests.get(movies_url)
z = zipfile.ZipFile(io.BytesIO(r.content))
movies_df = pd.read_csv(z.open('ml-latest-small/movies.csv'))
movie_lookup = dict(zip(movies_df['movieId'], movies_df['title']))
recommended_titles = [movie_lookup.get(mid, "Unknown") for mid in top_movies.index]
print("Recommended movie titles:", recommended_titles)
Recommended movie titles: ['Pursuit of Happyness, The (2006)', 'Eyes Wide Shut (1999)', 'Intouchables (2011)', 'Prestige, The (2006)', 'Limitless (2011)']

Your Turn: Experiment and Share#

  • Try changing parameters, or recommend for different users.
  • Can you come up with other ways to measure similarity? Try using correlation.
  • Share your results or code improvements in the comments if you are on YouTube!
  • Practice makes you more confident and creative.
# Advanced: Try item based collaborative filtering
item_similarity = cosine_similarity(rating_matrix.fillna(0).T)
print(item_similarity.shape)
 
(9724, 9724)

Wrap Up: Key Points from This Lesson#

  • Recommendation systems use massive data to personalize your experience.
  • Collaborative filtering suggests items based on patterns from many users.
  • The MovieLens dataset is a hands on, practical way to practice recommender logic in Python.
  • There is much more to learn, but you now know the basics needed to build your own suggestions!

Subscribe for More Python and Data Lessons!#

  • Like this lesson?
  • Subscribe to our YouTube channel and enable notifications.
  • You will learn new skills every week.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.