Lesson 14 · Data Mining
Understanding Association Rule Mining with the Apriori Algorithm: A Step-by-Step Guide
Ready to discover how supermarkets like Walmart figure out which items are usually bought together? In this lesson, we will explore the magic behind…
- CourseData Mining
- Lesson14 of 31
- Video17 min
- FormatJupyter notebook · 14 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbWeek 4: Association Rule Mining with Apriori Algorithm#
Ready to discover how supermarkets like Walmart figure out which items are usually bought together? In this lesson, we will explore the magic behind association rule mining using real-world grocery transactions.
You will learn how to load, prepare, and analyze shopping data step by step.
Our big goal: Find patterns in customer baskets and make smarter retail decisions.
# Suppress warnings for a clean output
import warnings
import numpy as np
np.random.seed(42)
warnings.filterwarnings("ignore")
Step 1: Load the Groceries Dataset#
We will work with a real-world groceries transactions dataset. Each row shows a single product that someone bought on a shopping trip.
Let us set up the data so we are ready for mining associations!
# Data setup (Groceries Dataset)
import pandas as pd
import kagglehub, os
path = kagglehub.dataset_download('heeraldedhia/groceries-dataset')
files = os.listdir(path)
csv_file = [f for f in files if f.endswith('.csv')][0]
df = pd.read_csv(os.path.join(path, csv_file))
print(df.shape)
print(df.head(3))
# Check for missing values in the data
df.isnull().sum()
# How many unique shoppers are in the data?
n_shoppers = df['Member_number'].nunique()
print(f"Unique shoppers: {n_shoppers}")
Step 2: Group Items into Baskets#
For association rule mining, we need to group all items each shopper bought into a single list, or 'basket'. This lets us look for patterns inside each shopping trip!
# Create transaction baskets: one list per Member_number
transactions = df.groupby(['Member_number'])['itemDescription'].apply(list).values.tolist()
print('First basket:', transactions[0])
# Count how many baskets (transactions) there are
print('Number of shopping baskets:', len(transactions))
Step 3: Prepare the Data for Apriori Algorithm#
The Apriori algorithm needs a special input format: rows as baskets, columns as items, with True or False for each item.
Let us convert our data to this 'one-hot encoded' table.
# Create a one-hot encoded DataFrame for market basket analysis
from mlxtend.preprocessing import TransactionEncoder
te = TransactionEncoder()
te_ary = te.fit(transactions).transform(transactions)
basket_df = pd.DataFrame(te_ary, columns=te.columns_)
print(basket_df.shape)
basket_df.head(3)
# What are the most common items in all baskets?
item_counts = basket_df.sum().sort_values(ascending=False)
print(item_counts.head(10))
# Step 4: Use Apriori to find frequent itemsets
from mlxtend.frequent_patterns import apriori
frequent_itemsets = apriori(basket_df, min_support=0.02, use_colnames=True)
print(frequent_itemsets.head())
What is 'Support' in Association Rule Mining?#
'Support' is how often an item or group shows up in all baskets.
For example, if milk is in 50 out of 100 baskets, its support is 0.5.
Higher support means more customers buy those items.
# Step 5: Generate association rules from the frequent itemsets
from mlxtend.frequent_patterns import association_rules
rules = association_rules(frequent_itemsets, metric='lift', min_threshold=1.1)
print(rules[['antecedents', 'consequents', 'support', 'confidence', 'lift']].head())
# Display rules with highest confidence
rules = rules.sort_values('confidence', ascending=False)
print(rules[['antecedents', 'consequents', 'support', 'confidence', 'lift']].head())
Step 6: Try Making Predictions#
Let us imagine being the store: Someone puts 'whole milk' in their basket. What would we suggest to them?
We will look for rules matching 'whole milk' to see what usually is bought next.
# Show rules where whole milk is in the basket
milk_rules = rules[rules['antecedents'].apply(lambda x: 'whole milk' in x)]
print(milk_rules[['antecedents', 'consequents', 'confidence', 'lift']].head())
# Optional: Try your own prediction!
item = input("Enter an item to see common associations: ")
filtered_rules = rules[rules['antecedents'].apply(lambda x: item in x)]
print(filtered_rules[['antecedents', 'consequents', 'confidence', 'lift']].head())
Recap and Next Steps#
This week, you loaded real transaction data, created baskets, and found patterns using the Apriori algorithm.
Association rule mining is a core tool for marketing, sales, and even website recommendations! Next, try changing the support and confidence thresholds to see how rules change.
Practice finding patterns in other datasetswhat rules might you discover in your favorite store?
# Challenge: Find the strongest rule for any item
challenge_item = input("Choose an item to investigate: ")
challenge_rules = rules[rules['antecedents'].apply(lambda x: challenge_item in x)]
if not challenge_rules.empty:
best = challenge_rules.sort_values('confidence', ascending=False).iloc[0]
print(f"If someone buys {challenge_item}, they often also buy {list(best['consequents'])[0]}. Confidence: {best['confidence']:.2f}")
else:
print("No strong rule found for this item. Try another!")
Thank You for Joining Week 4!#
Data mining helps reveal patterns in everyday life.
If you enjoyed learning, make sure to like and subscribe for more hands-on projects.
See you next time for new data discoveries!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



