Lesson 3 · Real-World Data Analytics
Python Data Analytics #03: Market Basket Analysis & Association Rules in Python
Video three of the hundred-video real-world data analytics series. Real support, confidence, and lift, built from scratch, on real invoices from the same…
- CourseReal-World Data Analytics
- Lesson3 of 26
- Video29 min
- FormatJupyter notebook · 30 code cells
- Data1 dataset
What you'll learn
Datasets used in this lesson
Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.
- online_retail_clean.csv36.2 MB
📓 Full notebook
Download .ipynbData Analytics 100, Video 3: Market Basket Analysis and Association Rules#
- Video three of the hundred-video real-world data analytics series.
- Real support, confidence, and lift, built from scratch, on real invoices from the same cleaned retail data.
- Let's get into it.
Part 1: What Market Basket Analysis Actually Measures#
import pandas as pd
from itertools import combinations
from collections import Counter
clean = pd.read_csv('online_retail_clean.csv', parse_dates=['InvoiceDate'])
clean.shape
Part 2: Restricting to Real UK Invoices#
uk = clean[clean['Country'] == 'United Kingdom']
uk.shape[0]
uk['InvoiceNo'].nunique()
Part 3: Real Top Products Worth Analyzing#
top_products = uk.groupby('Description')['InvoiceNo'].nunique().sort_values(ascending=False).head(40)
top_products.head(10)
top_set = set(top_products.index)
Part 4: Building the Real Baskets#
in_top = uk[uk['Description'].isin(top_set)]
baskets = in_top.groupby('InvoiceNo')['Description'].apply(set)
len(baskets)
Part 5: Keeping Only Real Multi-Item Baskets#
baskets = baskets[baskets.apply(len) >= 2]
len(baskets)
Part 6: Real Item Support#
n_invoices = uk['InvoiceNo'].nunique()
item_counts = Counter()
for basket in baskets:
for item in basket:
item_counts[item] += 1
support = {item: cnt / n_invoices for item, cnt in item_counts.items()}
sorted(support.items(), key=lambda x: -x[1])[:5]
Part 7: Real Pair Co-Occurrence Counts#
pair_counts = Counter()
for basket in baskets:
for a, b in combinations(sorted(basket), 2):
pair_counts[(a, b)] += 1
len(pair_counts)
Part 8: Real Pair Support#
pair_support = {pair: cnt / n_invoices for pair, cnt in pair_counts.items()}
sorted(pair_support.items(), key=lambda x: -x[1])[:5]
Part 9: Real Confidence, Both Directions#
def confidence(a, b):
key = tuple(sorted((a, b)))
return pair_counts[key] / item_counts[a]
example_a, example_b = sorted(pair_counts, key=lambda p: -pair_counts[p])[0]
round(confidence(example_a, example_b), 3), round(confidence(example_b, example_a), 3)
Part 10: Real Lift#
def lift(a, b):
key = tuple(sorted((a, b)))
return pair_support[key] / (support[a] * support[b])
round(lift(example_a, example_b), 2)
Part 11: Assembling One Real Rules Table#
rows = []
for (a, b), cnt in pair_counts.items():
rows.append((a, b, cnt, pair_support[(a, b)], confidence(a, b), confidence(b, a), lift(a, b)))
rules = pd.DataFrame(rows, columns=['A', 'B', 'co_count', 'support', 'conf_A_to_B', 'conf_B_to_A', 'lift'])
rules.shape[0]
Part 12: Real Top Rules by Lift#
rules.sort_values('lift', ascending=False).head(10)
Part 13: Real Top Rules by Confidence#
rules.sort_values('conf_A_to_B', ascending=False).head(10)[['A', 'B', 'conf_A_to_B', 'lift']]
Part 14: Filtering Out Real Rare Coincidences#
meaningful = rules[rules['co_count'] >= 20]
meaningful.shape[0]
meaningful.sort_values('lift', ascending=False).head(10)
Part 15: Interpreting One Real Rule in Plain Language#
top_rule = meaningful.sort_values('lift', ascending=False).iloc[0]
print(f"Customers who buy '{top_rule['A']}' are {round(top_rule['lift'], 1)}x more likely to also buy '{top_rule['B']}' than random chance would predict.")
Part 16: Real Average Lift as a Baseline#
rules['lift'].mean().round(2)
(rules['lift'] > 2).sum()
Part 17: Visualizing the Real Strongest Associations#
import matplotlib.pyplot as plt
top15 = meaningful.sort_values('lift', ascending=False).head(15)
labels = [f"{a[:15]}.. + {b[:15]}.." for a, b in zip(top15['A'], top15['B'])]
plt.figure(figsize=(9, 7))
plt.barh(labels[::-1], top15['lift'][::-1], color='darkorange')
plt.xlabel('Real Lift')
plt.title('Real Top 15 Real Product Associations by Lift')
plt.tight_layout()
plt.savefig('market_basket_top_lift.png', dpi=120)
plt.close()
Part 18: Real Support vs Real Lift#
plt.figure(figsize=(8, 6))
plt.scatter(rules['support'], rules['lift'], alpha=0.5, color='teal')
plt.xlabel('Real Support')
plt.ylabel('Real Lift')
plt.title('Real Support vs Real Lift Across All Candidate Pairs')
plt.tight_layout()
plt.savefig('market_basket_support_vs_lift.png', dpi=120)
plt.close()
Part 19: Which Real Product Shows Up in the Most Strong Rules#
strong = meaningful[meaningful['lift'] > 2]
appearance_counts = pd.concat([strong['A'], strong['B']]).value_counts()
appearance_counts.head(10)
Part 20: A Real Bundle Recommendation#
most_connected = appearance_counts.index[0]
bundle_candidates = strong[(strong['A'] == most_connected) | (strong['B'] == most_connected)]
bundle_candidates[['A', 'B', 'lift']].sort_values('lift', ascending=False)
Part 21: Saving the Real Rules Table#
meaningful.sort_values('lift', ascending=False).to_csv('market_basket_rules.csv', index=False)
reloaded_rules = pd.read_csv('market_basket_rules.csv')
reloaded_rules.shape[0] == meaningful.shape[0]
Part 22: Real Basket Size Distribution#
basket_sizes = baskets.apply(len)
basket_sizes.value_counts().sort_index()
basket_sizes.mean().round(2)
Part 23: Real Observed Pairs vs Real Possible Pairs#
from math import comb
possible_pairs = comb(len(top_set), 2)
len(pair_counts), possible_pairs
round(len(pair_counts) / possible_pairs * 100, 1)
Part 24: Manually Verifying One Real Support Value#
manual_count = sum(1 for basket in baskets if example_a in basket and example_b in basket)
manual_count == pair_counts[tuple(sorted((example_a, example_b)))]
Part 25: Why Confidence Is Not Symmetric#
item_counts[example_a], item_counts[example_b]
print(f"'{example_a}' appears in {item_counts[example_a]} real baskets; '{example_b}' appears in {item_counts[example_b]}, which is exactly why the two real confidence values differ.")
Part 26: Real Lift Value Distribution#
plt.figure(figsize=(8, 5))
plt.hist(rules['lift'], bins=30, color='slateblue', edgecolor='white')
plt.axvline(1, color='red', linestyle='--', label='Real independence (lift = 1)')
plt.xlabel('Real Lift')
plt.ylabel('Real Number of Rules')
plt.title('Real Distribution of Lift Across All Candidate Pairs')
plt.legend()
plt.tight_layout()
plt.savefig('market_basket_lift_distribution.png', dpi=120)
plt.close()
Part 27: Saving the Real Item Support Table#
support_df = pd.DataFrame(sorted(support.items(), key=lambda x: -x[1]), columns=['Product', 'Support'])
support_df.to_csv('market_basket_item_support.csv', index=False)
support_df.head(5)
Part 27b: Real Bundle Coverage Check#
baskets_with_strong_rule = sum(1 for basket in baskets if any(a in basket and b in basket for a, b in zip(strong['A'], strong['B'])))
round(baskets_with_strong_rule / len(baskets) * 100, 1)
Part 27c: Real Rule Count by Minimum Support Threshold#
for min_count in [5, 10, 20, 50, 100]:
n_rules = (rules['co_count'] >= min_count).sum()
print(f'min_count={min_count}: {n_rules} real rules survive')
Part 22: One Last Real Sanity Check#
(rules['support'] <= rules[['A']].join(pd.Series(support, name='sA'), on='A')['sA']).all()
Wrap-Up: What You Learned#
- Support, confidence, and lift are genuinely just three plain formulas over real invoice co-occurrence counts, no library required.
- Support measures how common a real product or pair is; confidence measures a real one-directional buying pattern; lift measures whether that pattern actually beats real chance.
- A real minimum support threshold is essential; rare pairs can show a huge lift purely by coincidence.
- The real most cross-sell-connected products are the strongest real candidates for bundle placement across the whole catalog.
- For production scale across a full real catalog, the mlxtend library implements this real same apriori logic far more efficiently; the math underneath is identical to what you just built.
- Next video: real cohort analysis, tracking how real customer groups actually retain over time.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



