Mathew K Analytics

Lesson 3 · Real-World Data Analytics

Python Data Analytics #03: Market Basket Analysis & Association Rules in Python

Video three of the hundred-video real-world data analytics series. Real support, confidence, and lift, built from scratch, on real invoices from the same…

What you'll learn

Datasets used in this lesson

Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.

📓 Full notebook

Download .ipynb

Data Analytics 100, Video 3: Market Basket Analysis and Association Rules#

  • Video three of the hundred-video real-world data analytics series.
  • Real support, confidence, and lift, built from scratch, on real invoices from the same cleaned retail data.
  • Let's get into it.

Part 1: What Market Basket Analysis Actually Measures#

import pandas as pd
from itertools import combinations
from collections import Counter
clean = pd.read_csv('online_retail_clean.csv', parse_dates=['InvoiceDate'])
clean.shape
(391150, 9)

Part 2: Restricting to Real UK Invoices#

uk = clean[clean['Country'] == 'United Kingdom']
uk.shape[0]
uk['InvoiceNo'].nunique()
16579

Part 3: Real Top Products Worth Analyzing#

top_products = uk.groupby('Description')['InvoiceNo'].nunique().sort_values(ascending=False).head(40)
top_products.head(10)
top_set = set(top_products.index)

Part 4: Building the Real Baskets#

in_top = uk[uk['Description'].isin(top_set)]
baskets = in_top.groupby('InvoiceNo')['Description'].apply(set)
len(baskets)
11365

Part 5: Keeping Only Real Multi-Item Baskets#

baskets = baskets[baskets.apply(len) >= 2]
len(baskets)
7716

Part 6: Real Item Support#

n_invoices = uk['InvoiceNo'].nunique()
item_counts = Counter()
for basket in baskets:
    for item in basket:
        item_counts[item] += 1
support = {item: cnt / n_invoices for item, cnt in item_counts.items()}
sorted(support.items(), key=lambda x: -x[1])[:5]
[('WHITE HANGING HEART T-LIGHT HOLDER', 0.09602509198383498),
 ('JUMBO BAG RED RETROSPOT', 0.0799806984739731),
 ('ASSORTED COLOUR BIRD ORNAMENT', 0.06640931298630798),
 ('PARTY BUNTING', 0.06634899571747391),
 ('LUNCH BAG RED RETROSPOT', 0.06598709210446951)]

Part 7: Real Pair Co-Occurrence Counts#

pair_counts = Counter()
for basket in baskets:
    for a, b in combinations(sorted(basket), 2):
        pair_counts[(a, b)] += 1
len(pair_counts)
780

Part 8: Real Pair Support#

pair_support = {pair: cnt / n_invoices for pair, cnt in pair_counts.items()}
sorted(pair_support.items(), key=lambda x: -x[1])[:5]
[(('JUMBO BAG PINK POLKADOT', 'JUMBO BAG RED RETROSPOT'), 0.030520538030038),
 (('LUNCH BAG  BLACK SKULL.', 'LUNCH BAG RED RETROSPOT'),
  0.029193558115688523),
 (('LUNCH BAG PINK POLKADOT', 'LUNCH BAG RED RETROSPOT'),
  0.028409433620845647),
 (('WOODEN FRAME ANTIQUE WHITE ', 'WOODEN PICTURE FRAME WHITE FINISH'),
  0.027625309126002775),
 (('ALARM CLOCK BAKELIKE GREEN', 'ALARM CLOCK BAKELIKE RED '),
  0.027384040050666504)]

Part 9: Real Confidence, Both Directions#

def confidence(a, b):
    key = tuple(sorted((a, b)))
    return pair_counts[key] / item_counts[a]
example_a, example_b = sorted(pair_counts, key=lambda p: -pair_counts[p])[0]
round(confidence(example_a, example_b), 3), round(confidence(example_b, example_a), 3)
(0.663, 0.382)

Part 10: Real Lift#

def lift(a, b):
    key = tuple(sorted((a, b)))
    return pair_support[key] / (support[a] * support[b])
round(lift(example_a, example_b), 2)
8.29

Part 11: Assembling One Real Rules Table#

rows = []
for (a, b), cnt in pair_counts.items():
    rows.append((a, b, cnt, pair_support[(a, b)], confidence(a, b), confidence(b, a), lift(a, b)))
rules = pd.DataFrame(rows, columns=['A', 'B', 'co_count', 'support', 'conf_A_to_B', 'conf_B_to_A', 'lift'])
rules.shape[0]
780

Part 12: Real Top Rules by Lift#

rules.sort_values('lift', ascending=False).head(10)
A B co_count support conf_A_to_B conf_B_to_A lift
22 ALARM CLOCK BAKELIKE GREEN ALARM CLOCK BAKELIKE RED 454 0.027384 0.702786 0.643059 16.503534
2 WOODEN FRAME ANTIQUE WHITE WOODEN PICTURE FRAME WHITE FINISH 458 0.027625 0.634349 0.564735 12.967784
13 HEART OF WICKER LARGE HEART OF WICKER SMALL 398 0.024006 0.537838 0.480097 10.756108
16 JAM MAKING SET PRINTED JAM MAKING SET WITH JARS 259 0.015622 0.399691 0.371593 9.507149
565 LUNCH BAG APPLE DESIGN LUNCH BAG SUKI DESIGN 329 0.019844 0.460140 0.402200 9.325989
646 JUMBO BAG ALPHABET JUMBO BAG VINTAGE LEAF 241 0.014536 0.374224 0.358631 9.232520
258 LUNCH BAG CARS BLUE LUNCH BAG PINK POLKADOT 396 0.023886 0.463700 0.474251 9.206810
165 LUNCH BAG BLACK SKULL. LUNCH BAG PINK POLKADOT 442 0.026660 0.460897 0.529341 9.151147
523 LUNCH BAG CARS BLUE LUNCH BAG SUKI DESIGN 363 0.021895 0.425059 0.443765 8.614970
31 LUNCH BAG PINK POLKADOT LUNCH BAG RED RETROSPOT 471 0.028409 0.564072 0.430530 8.548215

Part 13: Real Top Rules by Confidence#

rules.sort_values('conf_A_to_B', ascending=False).head(10)[['A', 'B', 'conf_A_to_B', 'lift']]
A B conf_A_to_B lift
22 ALARM CLOCK BAKELIKE GREEN ALARM CLOCK BAKELIKE RED 0.702786 16.503534
82 JUMBO BAG PINK POLKADOT JUMBO BAG RED RETROSPOT 0.663172 8.291647
2 WOODEN FRAME ANTIQUE WHITE WOODEN PICTURE FRAME WHITE FINISH 0.634349 12.967784
31 LUNCH BAG PINK POLKADOT LUNCH BAG RED RETROSPOT 0.564072 8.548215
13 HEART OF WICKER LARGE HEART OF WICKER SMALL 0.537838 10.756108
50 LUNCH BAG BLACK SKULL. LUNCH BAG RED RETROSPOT 0.504692 7.648350
54 LUNCH BAG CARS BLUE LUNCH BAG RED RETROSPOT 0.483607 7.328805
563 LUNCH BAG APPLE DESIGN LUNCH BAG RED RETROSPOT 0.465734 7.057960
258 LUNCH BAG CARS BLUE LUNCH BAG PINK POLKADOT 0.463700 9.206810
165 LUNCH BAG BLACK SKULL. LUNCH BAG PINK POLKADOT 0.460897 9.151147

Part 14: Filtering Out Real Rare Coincidences#

meaningful = rules[rules['co_count'] >= 20]
meaningful.shape[0]
meaningful.sort_values('lift', ascending=False).head(10)
A B co_count support conf_A_to_B conf_B_to_A lift
22 ALARM CLOCK BAKELIKE GREEN ALARM CLOCK BAKELIKE RED 454 0.027384 0.702786 0.643059 16.503534
2 WOODEN FRAME ANTIQUE WHITE WOODEN PICTURE FRAME WHITE FINISH 458 0.027625 0.634349 0.564735 12.967784
13 HEART OF WICKER LARGE HEART OF WICKER SMALL 398 0.024006 0.537838 0.480097 10.756108
16 JAM MAKING SET PRINTED JAM MAKING SET WITH JARS 259 0.015622 0.399691 0.371593 9.507149
565 LUNCH BAG APPLE DESIGN LUNCH BAG SUKI DESIGN 329 0.019844 0.460140 0.402200 9.325989
646 JUMBO BAG ALPHABET JUMBO BAG VINTAGE LEAF 241 0.014536 0.374224 0.358631 9.232520
258 LUNCH BAG CARS BLUE LUNCH BAG PINK POLKADOT 396 0.023886 0.463700 0.474251 9.206810
165 LUNCH BAG BLACK SKULL. LUNCH BAG PINK POLKADOT 442 0.026660 0.460897 0.529341 9.151147
523 LUNCH BAG CARS BLUE LUNCH BAG SUKI DESIGN 363 0.021895 0.425059 0.443765 8.614970
31 LUNCH BAG PINK POLKADOT LUNCH BAG RED RETROSPOT 471 0.028409 0.564072 0.430530 8.548215

Part 15: Interpreting One Real Rule in Plain Language#

top_rule = meaningful.sort_values('lift', ascending=False).iloc[0]
print(f"Customers who buy '{top_rule['A']}' are {round(top_rule['lift'], 1)}x more likely to also buy '{top_rule['B']}' than random chance would predict.")
Customers who buy 'ALARM CLOCK BAKELIKE GREEN' are 16.5x more likely to also buy 'ALARM CLOCK BAKELIKE RED ' than random chance would predict.

Part 16: Real Average Lift as a Baseline#

rules['lift'].mean().round(2)
(rules['lift'] > 2).sum()
np.int64(426)

Part 17: Visualizing the Real Strongest Associations#

import matplotlib.pyplot as plt
top15 = meaningful.sort_values('lift', ascending=False).head(15)
labels = [f"{a[:15]}.. + {b[:15]}.." for a, b in zip(top15['A'], top15['B'])]
plt.figure(figsize=(9, 7))
plt.barh(labels[::-1], top15['lift'][::-1], color='darkorange')
plt.xlabel('Real Lift')
plt.title('Real Top 15 Real Product Associations by Lift')
plt.tight_layout()
plt.savefig('market_basket_top_lift.png', dpi=120)
plt.close()

Part 18: Real Support vs Real Lift#

plt.figure(figsize=(8, 6))
plt.scatter(rules['support'], rules['lift'], alpha=0.5, color='teal')
plt.xlabel('Real Support')
plt.ylabel('Real Lift')
plt.title('Real Support vs Real Lift Across All Candidate Pairs')
plt.tight_layout()
plt.savefig('market_basket_support_vs_lift.png', dpi=120)
plt.close()

Part 19: Which Real Product Shows Up in the Most Strong Rules#

strong = meaningful[meaningful['lift'] > 2]
appearance_counts = pd.concat([strong['A'], strong['B']]).value_counts()
appearance_counts.head(10)
SET/5 RED RETROSPOT LID GLASS BOWLS    33
GARDENERS KNEELING PAD KEEP CALM       33
PACK OF 72 RETROSPOT CAKE CASES        32
RECIPE BOX PANTRY YELLOW DESIGN        32
SPOTTY BUNTING                         30
SET OF 4 PANTRY JELLY MOULDS           30
RETROSPOT TEA SET CERAMIC 11 PC        26
LUNCH BAG SPACEBOY DESIGN              26
LUNCH BAG RED RETROSPOT                25
LUNCH BAG SUKI DESIGN                  24
Name: count, dtype: int64

Part 20: A Real Bundle Recommendation#

most_connected = appearance_counts.index[0]
bundle_candidates = strong[(strong['A'] == most_connected) | (strong['B'] == most_connected)]
bundle_candidates[['A', 'B', 'lift']].sort_values('lift', ascending=False)
A B lift
136 PACK OF 72 RETROSPOT CAKE CASES SET/5 RED RETROSPOT LID GLASS BOWLS 4.228510
497 SET OF 4 PANTRY JELLY MOULDS SET/5 RED RETROSPOT LID GLASS BOWLS 4.128760
156 RECIPE BOX PANTRY YELLOW DESIGN SET/5 RED RETROSPOT LID GLASS BOWLS 4.098417
430 BAKING SET 9 PIECE RETROSPOT SET/5 RED RETROSPOT LID GLASS BOWLS 3.704557
401 SET OF 3 CAKE TINS PANTRY DESIGN SET/5 RED RETROSPOT LID GLASS BOWLS 3.575882
341 LUNCH BAG RED RETROSPOT SET/5 RED RETROSPOT LID GLASS BOWLS 3.348923
406 JAM MAKING SET PRINTED SET/5 RED RETROSPOT LID GLASS BOWLS 3.172687
438 RETROSPOT TEA SET CERAMIC 11 PC SET/5 RED RETROSPOT LID GLASS BOWLS 3.106128
404 JAM MAKING SET WITH JARS SET/5 RED RETROSPOT LID GLASS BOWLS 3.100907
135 ALARM CLOCK BAKELIKE RED SET/5 RED RETROSPOT LID GLASS BOWLS 3.061377
195 LUNCH BAG SPACEBOY DESIGN SET/5 RED RETROSPOT LID GLASS BOWLS 3.004345
330 SET/5 RED RETROSPOT LID GLASS BOWLS VINTAGE SNAP CARDS 2.967757
691 GARDENERS KNEELING PAD KEEP CALM SET/5 RED RETROSPOT LID GLASS BOWLS 2.950200
134 ALARM CLOCK BAKELIKE GREEN SET/5 RED RETROSPOT LID GLASS BOWLS 2.896900
191 LUNCH BAG CARS BLUE SET/5 RED RETROSPOT LID GLASS BOWLS 2.870336
162 SET/5 RED RETROSPOT LID GLASS BOWLS WHITE HANGING HEART T-LIGHT HOLDER 2.781467
339 LUNCH BAG PINK POLKADOT SET/5 RED RETROSPOT LID GLASS BOWLS 2.651554
537 LUNCH BAG SUKI DESIGN SET/5 RED RETROSPOT LID GLASS BOWLS 2.642215
148 HEART OF WICKER LARGE SET/5 RED RETROSPOT LID GLASS BOWLS 2.600153
337 LUNCH BAG BLACK SKULL. SET/5 RED RETROSPOT LID GLASS BOWLS 2.583550
161 REX CASH+CARRY JUMBO SHOPPER SET/5 RED RETROSPOT LID GLASS BOWLS 2.559372
96 JUMBO BAG RED RETROSPOT SET/5 RED RETROSPOT LID GLASS BOWLS 2.544334
599 LUNCH BAG APPLE DESIGN SET/5 RED RETROSPOT LID GLASS BOWLS 2.506747
283 SET/5 RED RETROSPOT LID GLASS BOWLS WOODEN PICTURE FRAME WHITE FINISH 2.437519
97 JUMBO STORAGE BAG SUKI SET/5 RED RETROSPOT LID GLASS BOWLS 2.427185
151 HEART OF WICKER SMALL SET/5 RED RETROSPOT LID GLASS BOWLS 2.352799
244 NATURAL SLATE HEART CHALKBOARD SET/5 RED RETROSPOT LID GLASS BOWLS 2.301842
155 PAPER CHAIN KIT 50'S CHRISTMAS SET/5 RED RETROSPOT LID GLASS BOWLS 2.282670
641 SET/5 RED RETROSPOT LID GLASS BOWLS SPOTTY BUNTING 2.260142
203 SET/5 RED RETROSPOT LID GLASS BOWLS WOODEN FRAME ANTIQUE WHITE 2.226898
481 PARTY BUNTING SET/5 RED RETROSPOT LID GLASS BOWLS 2.132578
239 JUMBO BAG PINK POLKADOT SET/5 RED RETROSPOT LID GLASS BOWLS 2.072690
760 HOT WATER BOTTLE KEEP CALM SET/5 RED RETROSPOT LID GLASS BOWLS 2.050044

Part 21: Saving the Real Rules Table#

meaningful.sort_values('lift', ascending=False).to_csv('market_basket_rules.csv', index=False)
reloaded_rules = pd.read_csv('market_basket_rules.csv')
reloaded_rules.shape[0] == meaningful.shape[0]
True

Part 22: Real Basket Size Distribution#

basket_sizes = baskets.apply(len)
basket_sizes.value_counts().sort_index()
basket_sizes.mean().round(2)
np.float64(4.15)

Part 23: Real Observed Pairs vs Real Possible Pairs#

from math import comb
possible_pairs = comb(len(top_set), 2)
len(pair_counts), possible_pairs
round(len(pair_counts) / possible_pairs * 100, 1)
100.0

Part 24: Manually Verifying One Real Support Value#

manual_count = sum(1 for basket in baskets if example_a in basket and example_b in basket)
manual_count == pair_counts[tuple(sorted((example_a, example_b)))]
True

Part 25: Why Confidence Is Not Symmetric#

item_counts[example_a], item_counts[example_b]
print(f"'{example_a}' appears in {item_counts[example_a]} real baskets; '{example_b}' appears in {item_counts[example_b]}, which is exactly why the two real confidence values differ.")
'JUMBO BAG PINK POLKADOT' appears in 763 real baskets; 'JUMBO BAG RED RETROSPOT' appears in 1326, which is exactly why the two real confidence values differ.

Part 26: Real Lift Value Distribution#

plt.figure(figsize=(8, 5))
plt.hist(rules['lift'], bins=30, color='slateblue', edgecolor='white')
plt.axvline(1, color='red', linestyle='--', label='Real independence (lift = 1)')
plt.xlabel('Real Lift')
plt.ylabel('Real Number of Rules')
plt.title('Real Distribution of Lift Across All Candidate Pairs')
plt.legend()
plt.tight_layout()
plt.savefig('market_basket_lift_distribution.png', dpi=120)
plt.close()

Part 27: Saving the Real Item Support Table#

support_df = pd.DataFrame(sorted(support.items(), key=lambda x: -x[1]), columns=['Product', 'Support'])
support_df.to_csv('market_basket_item_support.csv', index=False)
support_df.head(5)
Product Support
0 WHITE HANGING HEART T-LIGHT HOLDER 0.096025
1 JUMBO BAG RED RETROSPOT 0.079981
2 ASSORTED COLOUR BIRD ORNAMENT 0.066409
3 PARTY BUNTING 0.066349
4 LUNCH BAG RED RETROSPOT 0.065987

Part 27b: Real Bundle Coverage Check#

baskets_with_strong_rule = sum(1 for basket in baskets if any(a in basket and b in basket for a, b in zip(strong['A'], strong['B'])))
round(baskets_with_strong_rule / len(baskets) * 100, 1)
92.9

Part 27c: Real Rule Count by Minimum Support Threshold#

for min_count in [5, 10, 20, 50, 100]:
    n_rules = (rules['co_count'] >= min_count).sum()
    print(f'min_count={min_count}: {n_rules} real rules survive')
min_count=5: 780 real rules survive
min_count=10: 780 real rules survive
min_count=20: 779 real rules survive
min_count=50: 643 real rules survive
min_count=100: 231 real rules survive

Part 22: One Last Real Sanity Check#

(rules['support'] <= rules[['A']].join(pd.Series(support, name='sA'), on='A')['sA']).all()
np.True_

Wrap-Up: What You Learned#

  • Support, confidence, and lift are genuinely just three plain formulas over real invoice co-occurrence counts, no library required.
  • Support measures how common a real product or pair is; confidence measures a real one-directional buying pattern; lift measures whether that pattern actually beats real chance.
  • A real minimum support threshold is essential; rare pairs can show a huge lift purely by coincidence.
  • The real most cross-sell-connected products are the strongest real candidates for bundle placement across the whole catalog.
  • For production scale across a full real catalog, the mlxtend library implements this real same apriori logic far more efficiently; the math underneath is identical to what you just built.
  • Next video: real cohort analysis, tracking how real customer groups actually retain over time.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.