Mathew K Analytics

Lesson 14 · Data Mining

Understanding Association Rule Mining with the Apriori Algorithm: A Step-by-Step Guide

Ready to discover how supermarkets like Walmart figure out which items are usually bought together? In this lesson, we will explore the magic behind…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Week 4: Association Rule Mining with Apriori Algorithm#

Ready to discover how supermarkets like Walmart figure out which items are usually bought together? In this lesson, we will explore the magic behind association rule mining using real-world grocery transactions.

You will learn how to load, prepare, and analyze shopping data step by step.

Our big goal: Find patterns in customer baskets and make smarter retail decisions.

# Suppress warnings for a clean output
import warnings
import numpy as np
np.random.seed(42)
warnings.filterwarnings("ignore")

Step 1: Load the Groceries Dataset#

We will work with a real-world groceries transactions dataset. Each row shows a single product that someone bought on a shopping trip.

Let us set up the data so we are ready for mining associations!

# Data setup (Groceries Dataset)
import pandas as pd
import kagglehub, os
path = kagglehub.dataset_download('heeraldedhia/groceries-dataset')
files = os.listdir(path)
csv_file = [f for f in files if f.endswith('.csv')][0]
df = pd.read_csv(os.path.join(path, csv_file))
print(df.shape)
print(df.head(3))
(38765, 3)
   Member_number        Date itemDescription
0           1808  21-07-2015  tropical fruit
1           2552  05-01-2015      whole milk
2           2300  19-09-2015       pip fruit
# Check for missing values in the data
df.isnull().sum()
Member_number      0
Date               0
itemDescription    0
dtype: int64
# How many unique shoppers are in the data?
n_shoppers = df['Member_number'].nunique()
print(f"Unique shoppers: {n_shoppers}")
Unique shoppers: 3898

Step 2: Group Items into Baskets#

For association rule mining, we need to group all items each shopper bought into a single list, or 'basket'. This lets us look for patterns inside each shopping trip!

# Create transaction baskets: one list per Member_number
transactions = df.groupby(['Member_number'])['itemDescription'].apply(list).values.tolist()
print('First basket:', transactions[0])
First basket: ['soda', 'canned beer', 'sausage', 'sausage', 'whole milk', 'whole milk', 'pickled vegetables', 'misc. beverages', 'semi-finished bread', 'hygiene articles', 'yogurt', 'pastry', 'salty snack']
# Count how many baskets (transactions) there are
print('Number of shopping baskets:', len(transactions))
Number of shopping baskets: 3898

Step 3: Prepare the Data for Apriori Algorithm#

The Apriori algorithm needs a special input format: rows as baskets, columns as items, with True or False for each item.

Let us convert our data to this 'one-hot encoded' table.

# Create a one-hot encoded DataFrame for market basket analysis
from mlxtend.preprocessing import TransactionEncoder
te = TransactionEncoder()
te_ary = te.fit(transactions).transform(transactions)
basket_df = pd.DataFrame(te_ary, columns=te.columns_)
print(basket_df.shape)
basket_df.head(3)
(3898, 167)
Instant food products UHT-milk abrasive cleaner artif. sweetener baby cosmetics bags baking powder bathroom cleaner beef berries ... turkey vinegar waffles whipped/sour cream whisky white bread white wine whole milk yogurt zwieback
0 False False False False False False False False False False ... False False False False False False False True True False
1 False False False False False False False False True False ... False False False True False True False True False False
2 False False False False False False False False False False ... False False False False False False False True False False

3 rows × 167 columns

# What are the most common items in all baskets?
item_counts = basket_df.sum().sort_values(ascending=False)
print(item_counts.head(10))
whole milk          1786
other vegetables    1468
rolls/buns          1363
soda                1222
yogurt              1103
tropical fruit       911
root vegetables      899
bottled water        833
sausage              803
citrus fruit         723
dtype: int64
# Step 4: Use Apriori to find frequent itemsets
from mlxtend.frequent_patterns import apriori
frequent_itemsets = apriori(basket_df, min_support=0.02, use_colnames=True)
print(frequent_itemsets.head())
    support         itemsets
0  0.078502       (UHT-milk)
1  0.031042  (baking powder)
2  0.119548           (beef)
3  0.079785        (berries)
4  0.062083      (beverages)

What is 'Support' in Association Rule Mining?#

'Support' is how often an item or group shows up in all baskets.

For example, if milk is in 50 out of 100 baskets, its support is 0.5.

Higher support means more customers buy those items.

# Step 5: Generate association rules from the frequent itemsets
from mlxtend.frequent_patterns import association_rules
rules = association_rules(frequent_itemsets, metric='lift', min_threshold=1.1)
print(rules[['antecedents', 'consequents', 'support', 'confidence', 'lift']].head())
          antecedents         consequents   support  confidence      lift
0          (UHT-milk)     (bottled water)  0.021293    0.271242  1.269268
1     (bottled water)          (UHT-milk)  0.021293    0.099640  1.269268
2          (UHT-milk)  (other vegetables)  0.038994    0.496732  1.318979
3  (other vegetables)          (UHT-milk)  0.038994    0.103542  1.318979
4          (UHT-milk)        (rolls/buns)  0.031042    0.395425  1.130863
# Display rules with highest confidence
rules = rules.sort_values('confidence', ascending=False)
print(rules[['antecedents', 'consequents', 'support', 'confidence', 'lift']].head())
                                    antecedents   consequents   support  \
2417  (yogurt, bottled water, other vegetables)  (whole milk)  0.022063   
858               (bottled beer, shopping bags)  (whole milk)  0.020010   
2527     (yogurt, other vegetables, rolls/buns)  (whole milk)  0.034377   
1244               (shopping bags, canned beer)  (whole milk)  0.022063   
2568           (soda, yogurt, other vegetables)  (whole milk)  0.027963   

      confidence      lift  
2417    0.682540  1.489664  
858     0.661017  1.442690  
2527    0.656863  1.433623  
1244    0.656489  1.432806  
2568    0.648810  1.416047  

Step 6: Try Making Predictions#

Let us imagine being the store: Someone puts 'whole milk' in their basket. What would we suggest to them?

We will look for rules matching 'whole milk' to see what usually is bought next.

# Show rules where whole milk is in the basket
milk_rules = rules[rules['antecedents'].apply(lambda x: 'whole milk' in x)]
print(milk_rules[['antecedents', 'consequents', 'confidence', 'lift']].head())
                              antecedents         consequents  confidence  \
2400    (soda, whole milk, bottled water)  (other vegetables)    0.551282   
2414  (whole milk, yogurt, bottled water)  (other vegetables)    0.547771   
742                (UHT-milk, whole milk)  (other vegetables)    0.544304   
2458    (whole milk, rolls/buns, sausage)  (other vegetables)    0.536842   
2595          (whole milk, soda, sausage)        (rolls/buns)    0.525641   

          lift  
2400  1.463827  
2414  1.454503  
742   1.445297  
2458  1.425484  
2595  1.503264  
# Optional: Try your own prediction!
item = input("Enter an item to see common associations: ")
filtered_rules = rules[rules['antecedents'].apply(lambda x: item in x)]
print(filtered_rules[['antecedents', 'consequents', 'confidence', 'lift']].head())
                                     antecedents   consequents  confidence  \
2417   (yogurt, bottled water, other vegetables)  (whole milk)    0.682540   
2527      (yogurt, other vegetables, rolls/buns)  (whole milk)    0.656863   
2568            (soda, yogurt, other vegetables)  (whole milk)    0.648810   
2583  (yogurt, other vegetables, tropical fruit)  (whole milk)    0.648438   
2611               (yogurt, rolls/buns, sausage)  (whole milk)    0.640288   

          lift  
2417  1.489664  
2527  1.433623  
2568  1.416047  
2583  1.415235  
2611  1.397448  

Recap and Next Steps#

This week, you loaded real transaction data, created baskets, and found patterns using the Apriori algorithm.

Association rule mining is a core tool for marketing, sales, and even website recommendations! Next, try changing the support and confidence thresholds to see how rules change.

Practice finding patterns in other datasetswhat rules might you discover in your favorite store?

# Challenge: Find the strongest rule for any item
challenge_item = input("Choose an item to investigate: ")
challenge_rules = rules[rules['antecedents'].apply(lambda x: challenge_item in x)]
if not challenge_rules.empty:
    best = challenge_rules.sort_values('confidence', ascending=False).iloc[0]
    print(f"If someone buys {challenge_item}, they often also buy {list(best['consequents'])[0]}. Confidence: {best['confidence']:.2f}")
else:
    print("No strong rule found for this item. Try another!")
    
If someone buys rolls/buns, they often also buy whole milk. Confidence: 0.66

Thank You for Joining Week 4!#

Data mining helps reveal patterns in everyday life.

If you enjoyed learning, make sure to like and subscribe for more hands-on projects.

See you next time for new data discoveries!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.