Lesson 43 · Python for Retail E-commerce Analytics
Product Recommendation Systems Training for E-commerce Analytics with Python
Learn how e-commerce platforms make product recommendations using real retail data. Recommendation systems boost sales by showing the right products to…
- CoursePython for Retail E-commerce Analytics
- Lesson43 of 43
- Video20 min
- FormatJupyter notebook · 19 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbProduct Recommendation Systems in Retail Analytics#
- Learn how e-commerce platforms make product recommendations using real retail data.
- Recommendation systems boost sales by showing the right products to customers, improving both marketing and inventory usage.
- You will build and evaluate simple and advanced product recommendation logic, learning which products customers are most likely to buy together.
- No specialized machine learning libraries are needed; all logic will be built step by step.
import pandas as pd
import numpy as np
import warnings
warnings.filterwarnings('ignore')
Key Retail Analytics Concepts#
- Retail datasets track transactions, orders, products, and customer details.
- Every sale records the product, customer ID, date, quantity, and price.
- Beginners often confuse order-level and item-level data, or miscalculate revenue by using quantity only.
- Working with the right aggregation and grouping is key for recommendation systems.
# DATA SETUP: Load Online Retail Transactions Dataset (UCI)
url = 'https://archive.ics.uci.edu/ml/machine-learning-databases/00502/online_retail_II.xlsx'
retail_df = pd.read_excel(url, sheet_name='Year 2010-2011')
retail_df['InvoiceDate'] = pd.to_datetime(retail_df['InvoiceDate'])
print(retail_df.shape)
print(retail_df.head(3))
# Clean data: Remove missing Customer ID and Quantity <= 0
retail_df = retail_df.dropna(subset=['Customer ID'])
retail_df = retail_df[retail_df['Quantity'] > 0]
print('Cleaned shape:', retail_df.shape)
# Beginner Example 1: Count number of unique products bought
num_products = retail_df['StockCode'].nunique()
print('Number of unique products bought:', num_products)
# Beginner Example 2: Find total purchases per product
prod_counts = retail_df.groupby('StockCode').size().sort_values(ascending=False)
print(prod_counts.head(5))
# Beginner Example 3: List top 5 customers by transaction count
top_customers = retail_df['Customer ID'].value_counts().head(5)
print(top_customers)
# Intermediate Example 1: Create Customer-Product matrix
customer_product = pd.pivot_table(retail_df, index='Customer ID', columns='StockCode', values='Quantity', aggfunc='sum', fill_value=0)
print(customer_product.shape)
customer_product.head(3)
# Intermediate Example 2: Recommend top-3 popular products to every customer
top3_products = prod_counts.head(3).index.tolist()
print('Top 3 recommended products for all:', top3_products)
# Intermediate Example 3: Recommend products based on last purchase
def recommend_followup(customer_id):
purchases = retail_df[retail_df['Customer ID'] == customer_id].sort_values('InvoiceDate')
if purchases.shape[0] == 0:
return top3_products
last_purchased = purchases.iloc[-1]['StockCode']
# Find others that bought the same product, what else did they buy?
co_buyers = retail_df[retail_df['StockCode'] == last_purchased]['Customer ID'].unique()
also_bought = retail_df[retail_df['Customer ID'].isin(co_buyers)]
suggestions = also_bought['StockCode'].value_counts().index.tolist()
return [sku for sku in suggestions if sku != last_purchased][:3]
sample_id = retail_df['Customer ID'].iloc[0]
print('Follow-up recommendations for customer', sample_id, ':', recommend_followup(sample_id))
# Advanced Example 1: Product-to-product similarity based on co-purchases
from sklearn.metrics.pairwise import cosine_similarity
product_matrix = customer_product.T
similarity = cosine_similarity(product_matrix)
product_similarity = pd.DataFrame(similarity, index=product_matrix.index, columns=product_matrix.index)
sample_product = product_matrix.index[0]
sim_scores = product_similarity[sample_product].sort_values(ascending=False)[1:6]
print('Most similar products to', sample_product, ':')
print(sim_scores)
# Advanced Example 2: Association rule-style recommendation - people who bought X also bought Y
baskets = retail_df.groupby(['Invoice'])['StockCode'].apply(list)
from collections import Counter
pairs = []
for items in baskets:
pairs += [(a, b) for idx, a in enumerate(items) for b in items[idx + 1:]]
pair_counts = Counter(pairs)
common_pairs = pair_counts.most_common(5)
print('Top 5 product pairs bought together:')
for p in common_pairs:
print(p[0][0], 'and', p[0][1], ':', p[1], 'times')
# Advanced Example 3: Personalized recommendations using basket similarity (Jaccard index)
def jaccard_recommendation(target_customer):
target_purchases = set(retail_df[retail_df['Customer ID'] == target_customer]['StockCode'])
customer_sets = retail_df.groupby('Customer ID')['StockCode'].apply(set)
similarities = customer_sets.apply(lambda x: len(target_purchases & x) / len(target_purchases | x) if len(target_purchases | x) > 0 else 0)
most_similar = similarities.drop(target_customer).idxmax()
recommendations = customer_sets[most_similar] - target_purchases
return list(recommendations)[:3]
test_customer = retail_df['Customer ID'].iloc[100]
print('Personalized recommendations for customer', test_customer, ':', jaccard_recommendation(test_customer))
# Error Handling Example 1: What if a customer has no purchases?
empty_recommend = recommend_followup('made-up-customer')
print('Recommendation for non-existent customer:', empty_recommend)
# Error Handling Example 2: Incorrect aggregation - using mean instead of sum for product popularity
mean_popularity = retail_df.groupby('StockCode')['Quantity'].mean().sort_values(ascending=False).head(5)
print('By mean quantity per purchase:')
print(mean_popularity)
# Error Handling Example 3: Missing values in product or customer columns
retail_df_missing = retail_df.copy()
retail_df_missing.loc[retail_df_missing.sample(frac=0.001, random_state=42).index, 'StockCode'] = np.nan
missing_count = retail_df_missing['StockCode'].isnull().sum()
print('Artificially added missing StockCode count:', missing_count)
Retail Analytics Best Practices#
- Segment customers based on purchase frequency or category preference.
- Use total revenue, not just quantity, for product performance analysis.
- Analyze which products are frequently bought together for cross-selling.
- Always validate insights using domain knowledge before deploying.
# Best Practices Example: Identify high-value customers (customer segmentation)
retail_df['Revenue'] = retail_df['Quantity'] * retail_df['Price']
customer_revenue = retail_df.groupby('Customer ID')['Revenue'].sum().sort_values(ascending=False)
print('Top 5 customers by total revenue:')
print(customer_revenue.head(5))
# Best Practices Example: Analyze category-wise product sales
retail_product_demo = retail_df.groupby('StockCode').agg({'Quantity':'sum', 'Revenue':'sum'}).sort_values('Revenue', ascending=False)
print(retail_product_demo.head(5))
# Best Practices Example: Time-based demand trend for a popular product
pop_sku = prod_counts.index[0]
trend = retail_df[retail_df['StockCode'] == pop_sku].groupby(retail_df['InvoiceDate'].dt.month)['Quantity'].sum()
print('Monthly demand for top product:', pop_sku)
print(trend)
# End-to-end Mini Project: From raw data to recommendation insight
def recommend_cross_sell(customer_id):
last_products = retail_df[retail_df['Customer ID'] == customer_id]['StockCode'].unique()
baskets = retail_df.groupby('Invoice')['StockCode'].apply(set)
cross_sell_counts = Counter()
for basket in baskets:
if any(prod in basket for prod in last_products):
cross_sell_counts.update(basket - set(last_products))
recommendations = [prod for prod, count in cross_sell_counts.most_common(3)]
return recommendations
test_id = retail_df['Customer ID'].sample(1, random_state=42).values[0]
print('Best cross-sell recommendations for customer', test_id, ':', recommend_cross_sell(test_id))
Well Done! Next Steps#
- Practice by tuning recommendation rules or adding your own.
- Try using product descriptions for better recommendations.
- Ready to go further? Search YouTube for 'Retail Analytics in Python' to see more tutorials.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



