Lesson 11 · Data Mining
Foundations of Exploratory Data Mining in Descriptive Analytics
Welcome to Week 4! This lesson introduces you to exploratory data mining in Python. Data mining helps you discover patterns and insights from real-world…
- CourseData Mining
- Lesson11 of 31
- Video25 min
- FormatJupyter notebook · 18 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbWeek 4: Introduction to Exploratory Data Mining#
Welcome to Week 4! This lesson introduces you to exploratory data mining in Python.
Data mining helps you discover patterns and insights from real-world data.
We will use the Groceries Dataset to practice data preprocessing, exploratory analysis, association rules, classification, clustering, and more.
By the end, you will be confident to explore, clean, and analyze new datasets yourself.
import warnings
warnings.filterwarnings('ignore') # suppress warnings
# Data setup (Groceries Dataset)
import pandas as pd
import kagglehub, os
import numpy as np
np.random.seed(42)
path = kagglehub.dataset_download('heeraldedhia/groceries-dataset')
files = os.listdir(path)
csv_file = [f for f in files if f.endswith('.csv')][0]
df = pd.read_csv(os.path.join(path, csv_file))
print(df.shape)
print(df.head(3))
Step 1: Understanding the Dataset#
Our Groceries Dataset contains items purchased by shoppers.
Each row records a person, the shopping date, and which specific item they bought.
To use this for data mining, we will group all items bought together on the same shopping trip.
Let us explore what information is available.
# See unique columns and types
print(df.dtypes)
print(df.columns)
# Check for missing values
print(df.isnull().sum())
# Drop any rows with missing data (if any)
df = df.dropna()
print(df.shape)
# How many unique shoppers and unique items are there?
print('Unique shoppers:', df['Member_number'].nunique())
print('Unique items:', df['itemDescription'].nunique())
# Busiest shopping dates
print(df['Date'].value_counts().head(5))
Step 2: Preparing Basket Data#
For many data mining tasks, we need to know every item bought in each shopping trip.
Let us group the data by shopper and date, and collect all items into a 'basket.'
This basket format will help us discover popular item groups and rules.
# Create baskets: one set per shopper
basket_df = df.groupby(['Member_number'])['itemDescription'].apply(list).reset_index()
basket_df['basket_size'] = basket_df['itemDescription'].apply(len)
print(basket_df.head(3))
# How big are most baskets?
print(basket_df['basket_size'].describe())
# What are the top 10 most common items?
all_items = df['itemDescription']
print(all_items.value_counts().head(10))
Step 3: Discovering Association Rules#
Association rule mining finds which items are bought togetherlike 'if milk, then bread'.
To do this, we build a basket matrix: rows are baskets, columns are items, and cells show if the item was bought.
We use the mlxtend package and apriori algorithm.
from mlxtend.preprocessing import TransactionEncoder
# Convert baskets to dummy matrix for apriori
te = TransactionEncoder()
te_ary = te.fit(basket_df['itemDescription']).transform(basket_df['itemDescription'])
basket_matrix = pd.DataFrame(te_ary, columns=te.columns_)
print(basket_matrix.shape)
print(basket_matrix.head(3))
from mlxtend.frequent_patterns import apriori, association_rules
# Find frequent itemsets
frequent_itemsets = apriori(basket_matrix, min_support=0.02, use_colnames=True)
rules = association_rules(frequent_itemsets, metric='lift', min_threshold=1.2)
print(rules[['antecedents', 'consequents', 'support', 'confidence', 'lift']].head(5))
# Visualize top 5 rules by confidence
import matplotlib.pyplot as plt
top_rules = rules.sort_values('confidence', ascending=False).head(5)
plt.barh(range(5), top_rules['confidence'], tick_label=[str(list(a)[0]) + '' + str(list(c)[0]) for a, c in zip(top_rules['antecedents'], top_rules['consequents'])])
plt.xlabel('Confidence')
plt.title('Top 5 Association Rules')
plt.show()
Step 4: Basic Classification Setup#
For machine learning, classification aims to predict a label.
For practice, let us say we want to predict if a shopper will buy 'whole milk' or not.
This is a simplified classification task we can try with the groceries data.
# Add label for whole milk
basket_matrix['buys_whole_milk'] = basket_matrix['whole milk']
X = basket_matrix.drop('buys_whole_milk', axis=1)
y = basket_matrix['buys_whole_milk']
# Split into train and test sets
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, random_state=42)
print(X_train.shape, X_test.shape)
# Train a simple classifier (Logistic Regression)
from sklearn.linear_model import LogisticRegression
lr = LogisticRegression(max_iter=200)
lr.fit(X_train, y_train)
print('Test accuracy:', lr.score(X_test, y_test))
Step 5: Clustering with KMeans#
Clustering groups baskets with similar purchase patterns.
This is useful for finding customer segments automatically.
Let us cluster a sample using the KMeans algorithm from scikit-learn.
We will cluster using a smaller set of item columns for speed.
# Pick top 10 frequent items for clustering
cols = all_items.value_counts().head(10).index.tolist()
sample_basket = basket_matrix[cols]
from sklearn.cluster import KMeans
kmeans = KMeans(n_clusters=3, random_state=42)
clusters = kmeans.fit_predict(sample_basket)
sample_basket['cluster'] = clusters
print(sample_basket.groupby('cluster').mean())
Step 6: Mini ProjectMarket Basket Insights#
Practice: Create your own rule or cluster analysis.
Try changing the item or support threshold and see what new patterns appear.
You can analyze a different top item, or use more or fewer clusters in KMeans.
Exploring parameters is the best way to learn!
# Try it: predict another item's purchase
item = input('Enter an item name to predict (e.g., soda): ')
if item in basket_matrix.columns:
from sklearn.linear_model import LogisticRegression
y2 = basket_matrix[item]
X2 = basket_matrix.drop([item, 'buys_whole_milk'], axis=1, errors='ignore')
X2_train, X2_test, y2_train, y2_test = train_test_split(X2, y2, test_size=0.25, random_state=42)
lr2 = LogisticRegression(max_iter=100)
lr2.fit(X2_train, y2_train)
print(f'Accuracy in predicting {item}:', lr2.score(X2_test, y2_test))
else:
print('Item not found. Try a different name.')
# Explore: how many clusters work best?
for k in range(2, 6):
kmeans_k = KMeans(n_clusters=k, random_state=42).fit(sample_basket[cols])
print(f'k={k}: Inertia={kmeans_k.inertia_:.2f}')
Troubleshooting and Best Practices#
If something does not work, check for typos and missing dependencies.
Always inspect your data before trusting results.
Test your logic with small samples to spot errors early.
Document your process in comments and markdown cells.
Keep your code tidy and backed up.
Extra Tips#
Try more datasets, like movies or sports, for more practice.
Many data mining concepts also work for text or images.
Explore the scikit-learn and pandas documentation for more features.
Sharing your findings with friends helps you learn faster!
Challenge Exercises#
Find a set of three items that often appear together.
Predict a different item's purchase, such as 'soda' or 'yogurt.'
Cluster with five or more clusters and describe what you observe.
Try visualizing the basket sizes with a histogram.
Share your solutions in the comments to help other learners!
Recap#
In this lesson, you learned:
How to explore, clean, and prepare transactional data. How to group transactions into baskets for mining. Basic association rule mining for shopping pattern discovery. Simple classification and clustering from real data. Why testing, visualization, and parameter searching matter.
Keep practicing with new data to build your data mining confidence!
Thank you for learning Exploratory Data Mining!#
If you enjoyed this lesson, please like, subscribe, and share.
Check out the next video for deeper dives into machine learning!
Happy mining!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



