Mathew K Analytics

Lesson 43 · Python for Retail E-commerce Analytics

Product Recommendation Systems Training for E-commerce Analytics with Python

Learn how e-commerce platforms make product recommendations using real retail data. Recommendation systems boost sales by showing the right products to…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Product Recommendation Systems in Retail Analytics#

  • Learn how e-commerce platforms make product recommendations using real retail data.
  • Recommendation systems boost sales by showing the right products to customers, improving both marketing and inventory usage.
  • You will build and evaluate simple and advanced product recommendation logic, learning which products customers are most likely to buy together.
  • No specialized machine learning libraries are needed; all logic will be built step by step.
import pandas as pd
import numpy as np
import warnings
warnings.filterwarnings('ignore')

Key Retail Analytics Concepts#

  • Retail datasets track transactions, orders, products, and customer details.
  • Every sale records the product, customer ID, date, quantity, and price.
  • Beginners often confuse order-level and item-level data, or miscalculate revenue by using quantity only.
  • Working with the right aggregation and grouping is key for recommendation systems.
# DATA SETUP: Load Online Retail Transactions Dataset (UCI)
url = 'https://archive.ics.uci.edu/ml/machine-learning-databases/00502/online_retail_II.xlsx'
retail_df = pd.read_excel(url, sheet_name='Year 2010-2011')
retail_df['InvoiceDate'] = pd.to_datetime(retail_df['InvoiceDate'])
print(retail_df.shape)
print(retail_df.head(3))
(541910, 8)
  Invoice StockCode                         Description  Quantity  \
0  536365    85123A  WHITE HANGING HEART T-LIGHT HOLDER         6   
1  536365     71053                 WHITE METAL LANTERN         6   
2  536365    84406B      CREAM CUPID HEARTS COAT HANGER         8   

          InvoiceDate  Price  Customer ID         Country  
0 2010-12-01 08:26:00   2.55      17850.0  United Kingdom  
1 2010-12-01 08:26:00   3.39      17850.0  United Kingdom  
2 2010-12-01 08:26:00   2.75      17850.0  United Kingdom  
# Clean data: Remove missing Customer ID and Quantity <= 0
retail_df = retail_df.dropna(subset=['Customer ID'])
retail_df = retail_df[retail_df['Quantity'] > 0]
print('Cleaned shape:', retail_df.shape)
Cleaned shape: (397925, 8)
# Beginner Example 1: Count number of unique products bought
num_products = retail_df['StockCode'].nunique()
print('Number of unique products bought:', num_products)
Number of unique products bought: 3665
# Beginner Example 2: Find total purchases per product
prod_counts = retail_df.groupby('StockCode').size().sort_values(ascending=False)
print(prod_counts.head(5))
StockCode
85123A    2035
22423     1724
85099B    1618
84879     1408
47566     1397
dtype: int64
# Beginner Example 3: List top 5 customers by transaction count
top_customers = retail_df['Customer ID'].value_counts().head(5)
print(top_customers)
Customer ID
17841.0    7847
14911.0    5677
14096.0    5111
12748.0    4596
14606.0    2700
Name: count, dtype: int64
# Intermediate Example 1: Create Customer-Product matrix
customer_product = pd.pivot_table(retail_df, index='Customer ID', columns='StockCode', values='Quantity', aggfunc='sum', fill_value=0)
print(customer_product.shape)
customer_product.head(3)
(4339, 3665)
StockCode 10002 10080 10120 10125 10133 10135 11001 15030 15034 15036 ... 90214V 90214W 90214Y 90214Z BANK CHARGES C2 DOT M PADS POST
Customer ID
12346.0 0 0 0 0 0 0 0 0 0 0 ... 0 0 0 0 0 0 0 0 0 0
12347.0 0 0 0 0 0 0 0 0 0 0 ... 0 0 0 0 0 0 0 0 0 0
12348.0 0 0 0 0 0 0 0 0 0 0 ... 0 0 0 0 0 0 0 0 0 9

3 rows × 3665 columns

# Intermediate Example 2: Recommend top-3 popular products to every customer
top3_products = prod_counts.head(3).index.tolist()
print('Top 3 recommended products for all:', top3_products)
Top 3 recommended products for all: ['85123A', 22423, '85099B']
# Intermediate Example 3: Recommend products based on last purchase
def recommend_followup(customer_id):
    purchases = retail_df[retail_df['Customer ID'] == customer_id].sort_values('InvoiceDate')
    if purchases.shape[0] == 0:
        return top3_products
    last_purchased = purchases.iloc[-1]['StockCode']
    # Find others that bought the same product, what else did they buy?
    co_buyers = retail_df[retail_df['StockCode'] == last_purchased]['Customer ID'].unique()
    also_bought = retail_df[retail_df['Customer ID'].isin(co_buyers)]
    suggestions = also_bought['StockCode'].value_counts().index.tolist()
    return [sku for sku in suggestions if sku != last_purchased][:3]

sample_id = retail_df['Customer ID'].iloc[0]
print('Follow-up recommendations for customer', sample_id, ':', recommend_followup(sample_id))
Follow-up recommendations for customer 17850.0 : ['85123A', 22865, '85099B']
# Advanced Example 1: Product-to-product similarity based on co-purchases
from sklearn.metrics.pairwise import cosine_similarity
product_matrix = customer_product.T
similarity = cosine_similarity(product_matrix)
product_similarity = pd.DataFrame(similarity, index=product_matrix.index, columns=product_matrix.index)
sample_product = product_matrix.index[0]
sim_scores = product_similarity[sample_product].sort_values(ascending=False)[1:6]
print('Most similar products to', sample_product, ':')
print(sim_scores)
Most similar products to 10002 :
StockCode
10125     0.853890
23224     0.713423
23222     0.699006
85014A    0.626077
20682     0.597314
Name: 10002, dtype: float64
# Advanced Example 2: Association rule-style recommendation - people who bought X also bought Y
baskets = retail_df.groupby(['Invoice'])['StockCode'].apply(list)
from collections import Counter
pairs = []
for items in baskets:
    pairs += [(a, b) for idx, a in enumerate(items) for b in items[idx + 1:]]
pair_counts = Counter(pairs)
common_pairs = pair_counts.most_common(5)
print('Top 5 product pairs bought together:')
for p in common_pairs:
    print(p[0][0], 'and', p[0][1], ':', p[1], 'times')
Top 5 product pairs bought together:
22697 and 22698 : 444 times
20725 and 22384 : 327 times
20725 and 20727 : 316 times
22386 and 85099B : 315 times
22699 and 22697 : 314 times
# Advanced Example 3: Personalized recommendations using basket similarity (Jaccard index)
def jaccard_recommendation(target_customer):
    target_purchases = set(retail_df[retail_df['Customer ID'] == target_customer]['StockCode'])
    customer_sets = retail_df.groupby('Customer ID')['StockCode'].apply(set)
    similarities = customer_sets.apply(lambda x: len(target_purchases & x) / len(target_purchases | x) if len(target_purchases | x) > 0 else 0)
    most_similar = similarities.drop(target_customer).idxmax()
    recommendations = customer_sets[most_similar] - target_purchases
    return list(recommendations)[:3]

test_customer = retail_df['Customer ID'].iloc[100]
print('Personalized recommendations for customer', test_customer, ':', jaccard_recommendation(test_customer))
Personalized recommendations for customer 14688.0 : [23582, 23200, 23202]
# Error Handling Example 1: What if a customer has no purchases?
empty_recommend = recommend_followup('made-up-customer')
print('Recommendation for non-existent customer:', empty_recommend)
Recommendation for non-existent customer: ['85123A', 22423, '85099B']
# Error Handling Example 2: Incorrect aggregation - using mean instead of sum for product popularity
mean_popularity = retail_df.groupby('StockCode')['Quantity'].mean().sort_values(ascending=False).head(5)
print('By mean quantity per purchase:')
print(mean_popularity)
By mean quantity per purchase:
StockCode
23843     80995.000000
47556B     1300.000000
84568       520.000000
23166       393.515152
84826       380.611111
Name: Quantity, dtype: float64
# Error Handling Example 3: Missing values in product or customer columns
retail_df_missing = retail_df.copy()
retail_df_missing.loc[retail_df_missing.sample(frac=0.001, random_state=42).index, 'StockCode'] = np.nan
missing_count = retail_df_missing['StockCode'].isnull().sum()
print('Artificially added missing StockCode count:', missing_count)
Artificially added missing StockCode count: 398

Retail Analytics Best Practices#

  • Segment customers based on purchase frequency or category preference.
  • Use total revenue, not just quantity, for product performance analysis.
  • Analyze which products are frequently bought together for cross-selling.
  • Always validate insights using domain knowledge before deploying.
# Best Practices Example: Identify high-value customers (customer segmentation)
retail_df['Revenue'] = retail_df['Quantity'] * retail_df['Price']
customer_revenue = retail_df.groupby('Customer ID')['Revenue'].sum().sort_values(ascending=False)
print('Top 5 customers by total revenue:')
print(customer_revenue.head(5))
Top 5 customers by total revenue:
Customer ID
14646.0    280206.02
18102.0    259657.30
17450.0    194550.79
16446.0    168472.50
14911.0    143825.06
Name: Revenue, dtype: float64
# Best Practices Example: Analyze category-wise product sales
retail_product_demo = retail_df.groupby('StockCode').agg({'Quantity':'sum', 'Revenue':'sum'}).sort_values('Revenue', ascending=False)
print(retail_product_demo.head(5))
           Quantity    Revenue
StockCode                     
23843         80995  168469.60
22423         12412  142592.95
85123A        36782  100603.50
85099B        46181   85220.78
23166         77916   81416.73
# Best Practices Example: Time-based demand trend for a popular product
pop_sku = prod_counts.index[0]
trend = retail_df[retail_df['StockCode'] == pop_sku].groupby(retail_df['InvoiceDate'].dt.month)['Quantity'].sum()
print('Monthly demand for top product:', pop_sku)
print(trend)
Monthly demand for top product: 85123A
InvoiceDate
1     5467
2     1823
3     1918
4     3725
5     3846
6     1618
7     2971
8     2046
9     2444
10    1650
11    4861
12    4413
Name: Quantity, dtype: int64
# End-to-end Mini Project: From raw data to recommendation insight
def recommend_cross_sell(customer_id):
    last_products = retail_df[retail_df['Customer ID'] == customer_id]['StockCode'].unique()
    baskets = retail_df.groupby('Invoice')['StockCode'].apply(set)
    cross_sell_counts = Counter()
    for basket in baskets:
        if any(prod in basket for prod in last_products):
            cross_sell_counts.update(basket - set(last_products))
    recommendations = [prod for prod, count in cross_sell_counts.most_common(3)]
    return recommendations
test_id = retail_df['Customer ID'].sample(1, random_state=42).values[0]
print('Best cross-sell recommendations for customer', test_id, ':', recommend_cross_sell(test_id))
Best cross-sell recommendations for customer 17894.0 : [22423, '85099B', 84879]

Well Done! Next Steps#

  • Practice by tuning recommendation rules or adding your own.
  • Try using product descriptions for better recommendations.
  • Ready to go further? Search YouTube for 'Retail Analytics in Python' to see more tutorials.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.