Lesson 6 · Pandas Projects
Restaurant Reviews & Ratings Analysis with Python Pandas
A complete, standalone tutorial: build a synthetic reviews dataset, clean out duplicates, then rank restaurants and cuisines with pandas. No prior pandas…
- CoursePandas Projects
- Lesson6 of 10
- Video16 min
- FormatJupyter notebook · 14 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbPandas for Restaurant Reviews: Ratings, Cuisines, and Duplicates#
- A complete, standalone tutorial: build a synthetic reviews dataset, clean out duplicates, then rank restaurants and cuisines with pandas.
- No prior pandas experience needed. Let's jump straight in.
Before You Start#
- Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
- If pandas isn't installed yet, open a terminal in VS Code and run: pip install pandas
Part 1: Building the Dataset#
import pandas as pd
import numpy as np
print(pd.__version__)
Building a Restaurant List#
rng = np.random.default_rng(seed=99)
cuisines = ['Italian', 'Japanese', 'Mexican', 'Indian', 'Thai']
cities = ['Portside', 'Elmfield', 'Brookhaven']
price_ranges = ['Budget', 'Mid-Range', 'Upscale']
restaurants = pd.DataFrame({
'restaurant_id': range(1, 26),
'restaurant_name': [f'Restaurant {i}' for i in range(1, 26)],
'cuisine': rng.choice(cuisines, size=25),
'city': rng.choice(cities, size=25),
'price_range': rng.choice(price_ranges, size=25, p=[0.35, 0.45, 0.2])
})
restaurants.head()
Generating Reviews#
n_reviews = 600
review_rows = []
review_dates = pd.date_range('2026-01-01', periods=210, freq='D')
for review_id in range(1, n_reviews + 1):
restaurant_id = int(rng.choice(restaurants['restaurant_id']))
rating = int(rng.choice([1, 2, 3, 4, 5], p=[0.05, 0.1, 0.2, 0.35, 0.3]))
review_date = rng.choice(review_dates)
review_rows.append([review_id, restaurant_id, rating, review_date])
reviews = pd.DataFrame(review_rows, columns=['review_id', 'restaurant_id', 'rating', 'review_date'])
print(reviews.shape)
Injecting Duplicate Reviews#
duplicate_rows = reviews.sample(n=20, random_state=5)
reviews = pd.concat([reviews, duplicate_rows], ignore_index=True)
reviews.to_csv('restaurant_reviews.csv', index=False)
reviews = pd.read_csv('restaurant_reviews.csv', parse_dates=['review_date'])
print(reviews.shape)
Part 2: First Look and Deduplication#
print(reviews.duplicated().sum())
reviews[reviews.duplicated()].head()
reviews = reviews.drop_duplicates().reset_index(drop=True)
print(reviews.shape)
print(reviews.duplicated().sum())
Part 3: Merging and Ranking Restaurants#
full_reviews = reviews.merge(restaurants, on='restaurant_id', how='left')
full_reviews[['restaurant_name', 'cuisine', 'city', 'rating']].head()
restaurant_stats = full_reviews.groupby('restaurant_name').agg(
review_count=('rating', 'count'),
avg_rating=('rating', 'mean')
).round(2)
top_rated = restaurant_stats[restaurant_stats['review_count'] >= 10].sort_values('avg_rating', ascending=False)
top_rated.head(5)
Part 4: Cuisine Analysis#
cuisine_summary = full_reviews.groupby('cuisine').agg(
review_count=('rating', 'count'),
avg_rating=('rating', 'mean')
).round(2).sort_values('avg_rating', ascending=False)
cuisine_summary
full_reviews['rating'].value_counts().sort_index()
Part 5: Pivot Table and Cross-Tab#
cuisine_price_pivot = pd.pivot_table(full_reviews, values='rating', index='cuisine', columns='price_range', aggfunc='mean').round(2)
cuisine_price_pivot = cuisine_price_pivot[['Budget', 'Mid-Range', 'Upscale']]
cuisine_price_pivot
pd.crosstab(full_reviews['cuisine'], full_reviews['city'])
Part 6: Visualizing the Results#
import matplotlib
matplotlib.use('Agg')
import matplotlib.pyplot as plt
fig, ax = plt.subplots(figsize=(8, 5))
cuisine_summary['avg_rating'].plot(kind='bar', ax=ax, color='darkorange', ylim=(0, 5))
ax.set_title('Average Rating by Cuisine')
ax.set_ylabel('Average Rating (out of 5)')
plt.tight_layout()
plt.savefig('avg_rating_by_cuisine.png', dpi=150)
plt.close(fig)
print('Saved avg_rating_by_cuisine.png')
fig, ax = plt.subplots(figsize=(8, 5))
cuisine_price_pivot.plot(kind='bar', ax=ax, ylim=(0, 5))
ax.set_title('Average Rating by Cuisine and Price Range')
ax.set_ylabel('Average Rating (out of 5)')
ax.legend(title='Price Range')
plt.tight_layout()
plt.savefig('rating_by_cuisine_and_price.png', dpi=150)
plt.close(fig)
print('Saved rating_by_cuisine_and_price.png')
Wrap-Up: What You Learned#
- Generating a realistic synthetic reviews dataset with deliberately injected duplicates, then saving and reloading with to_csv and read_csv.
- Detecting duplicate rows with duplicated, and removing them with drop_duplicates.
- Merging reviews onto restaurant details, then ranking restaurants with a minimum-review-count filter.
- Cuisine-level summaries with groupby, and a full rating distribution with value_counts.
- A cuisine-by-price pivot_table, and a cuisine-by-city crosstab.
- Grouped bar charts, including one plotted directly from a pivoted DataFrame, with matplotlib.
- You went from a duplicate-riddled synthetic reviews feed to a fully cleaned, ranked restaurant report. If you want the next dataset in this series to land in your feed automatically, subscribing is the move see you in the next one.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



