Mathew K Analytics

Lesson 11 · Data Mining

Foundations of Exploratory Data Mining in Descriptive Analytics

Welcome to Week 4! This lesson introduces you to exploratory data mining in Python. Data mining helps you discover patterns and insights from real-world…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Week 4: Introduction to Exploratory Data Mining#

Welcome to Week 4! This lesson introduces you to exploratory data mining in Python.

Data mining helps you discover patterns and insights from real-world data.

We will use the Groceries Dataset to practice data preprocessing, exploratory analysis, association rules, classification, clustering, and more.

By the end, you will be confident to explore, clean, and analyze new datasets yourself.

import warnings
warnings.filterwarnings('ignore') # suppress warnings

# Data setup (Groceries Dataset)
import pandas as pd
import kagglehub, os
import numpy as np
np.random.seed(42)
path = kagglehub.dataset_download('heeraldedhia/groceries-dataset')
files = os.listdir(path)
csv_file = [f for f in files if f.endswith('.csv')][0]
df = pd.read_csv(os.path.join(path, csv_file))
print(df.shape)
print(df.head(3))
(38765, 3)
   Member_number        Date itemDescription
0           1808  21-07-2015  tropical fruit
1           2552  05-01-2015      whole milk
2           2300  19-09-2015       pip fruit

Step 1: Understanding the Dataset#

Our Groceries Dataset contains items purchased by shoppers.

Each row records a person, the shopping date, and which specific item they bought.

To use this for data mining, we will group all items bought together on the same shopping trip.

Let us explore what information is available.

# See unique columns and types
print(df.dtypes)
print(df.columns)
Member_number       int64
Date               object
itemDescription    object
dtype: object
Index(['Member_number', 'Date', 'itemDescription'], dtype='object')
# Check for missing values
print(df.isnull().sum())
Member_number      0
Date               0
itemDescription    0
dtype: int64
# Drop any rows with missing data (if any)
df = df.dropna()
print(df.shape)
(38765, 3)
# How many unique shoppers and unique items are there?
print('Unique shoppers:', df['Member_number'].nunique())
print('Unique items:', df['itemDescription'].nunique())
Unique shoppers: 3898
Unique items: 167
# Busiest shopping dates
print(df['Date'].value_counts().head(5))
Date
21-01-2015    96
21-07-2015    93
08-08-2015    92
29-11-2015    92
30-04-2015    91
Name: count, dtype: int64

Step 2: Preparing Basket Data#

For many data mining tasks, we need to know every item bought in each shopping trip.

Let us group the data by shopper and date, and collect all items into a 'basket.'

This basket format will help us discover popular item groups and rules.

# Create baskets: one set per shopper
basket_df = df.groupby(['Member_number'])['itemDescription'].apply(list).reset_index()
basket_df['basket_size'] = basket_df['itemDescription'].apply(len)
print(basket_df.head(3))
   Member_number                                    itemDescription  \
0           1000  [soda, canned beer, sausage, sausage, whole mi...   
1           1001  [frankfurter, frankfurter, beef, sausage, whol...   
2           1002  [tropical fruit, butter milk, butter, frozen v...   

   basket_size  
0           13  
1           12  
2            8  
# How big are most baskets?
print(basket_df['basket_size'].describe())
count    3898.000000
mean        9.944844
std         5.310796
min         2.000000
25%         6.000000
50%         9.000000
75%        13.000000
max        36.000000
Name: basket_size, dtype: float64
# What are the top 10 most common items?
all_items = df['itemDescription']
print(all_items.value_counts().head(10))
itemDescription
whole milk          2502
other vegetables    1898
rolls/buns          1716
soda                1514
yogurt              1334
root vegetables     1071
tropical fruit      1032
bottled water        933
sausage              924
citrus fruit         812
Name: count, dtype: int64

Step 3: Discovering Association Rules#

Association rule mining finds which items are bought togetherlike 'if milk, then bread'.

To do this, we build a basket matrix: rows are baskets, columns are items, and cells show if the item was bought.

We use the mlxtend package and apriori algorithm.

from mlxtend.preprocessing import TransactionEncoder

# Convert baskets to dummy matrix for apriori
te = TransactionEncoder()
te_ary = te.fit(basket_df['itemDescription']).transform(basket_df['itemDescription'])
basket_matrix = pd.DataFrame(te_ary, columns=te.columns_)
print(basket_matrix.shape)
print(basket_matrix.head(3))
(3898, 167)
   Instant food products  UHT-milk  abrasive cleaner  artif. sweetener  \
0                  False     False             False             False   
1                  False     False             False             False   
2                  False     False             False             False   

   baby cosmetics   bags  baking powder  bathroom cleaner   beef  berries  \
0           False  False          False             False  False    False   
1           False  False          False             False   True    False   
2           False  False          False             False  False    False   

   ...  turkey  vinegar  waffles  whipped/sour cream  whisky  white bread  \
0  ...   False    False    False               False   False        False   
1  ...   False    False    False                True   False         True   
2  ...   False    False    False               False   False        False   

   white wine  whole milk  yogurt  zwieback  
0       False        True    True     False  
1       False        True   False     False  
2       False        True   False     False  

[3 rows x 167 columns]
from mlxtend.frequent_patterns import apriori, association_rules

# Find frequent itemsets
frequent_itemsets = apriori(basket_matrix, min_support=0.02, use_colnames=True)
rules = association_rules(frequent_itemsets, metric='lift', min_threshold=1.2)
print(rules[['antecedents', 'consequents', 'support', 'confidence', 'lift']].head(5))
          antecedents         consequents   support  confidence      lift
0     (bottled water)          (UHT-milk)  0.021293    0.099640  1.269268
1          (UHT-milk)     (bottled water)  0.021293    0.271242  1.269268
2  (other vegetables)          (UHT-milk)  0.038994    0.103542  1.318979
3          (UHT-milk)  (other vegetables)  0.038994    0.496732  1.318979
4              (beef)            (butter)  0.020523    0.171674  1.357372
# Visualize top 5 rules by confidence
import matplotlib.pyplot as plt

top_rules = rules.sort_values('confidence', ascending=False).head(5)
plt.barh(range(5), top_rules['confidence'], tick_label=[str(list(a)[0]) + '' + str(list(c)[0]) for a, c in zip(top_rules['antecedents'], top_rules['consequents'])])
plt.xlabel('Confidence')
plt.title('Top 5 Association Rules')
plt.show()
No description has been provided for this image

Step 4: Basic Classification Setup#

For machine learning, classification aims to predict a label.

For practice, let us say we want to predict if a shopper will buy 'whole milk' or not.

This is a simplified classification task we can try with the groceries data.

# Add label for whole milk
basket_matrix['buys_whole_milk'] = basket_matrix['whole milk']

X = basket_matrix.drop('buys_whole_milk', axis=1)
y = basket_matrix['buys_whole_milk']
# Split into train and test sets
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, random_state=42)
print(X_train.shape, X_test.shape)
(2923, 167) (975, 167)
# Train a simple classifier (Logistic Regression)
from sklearn.linear_model import LogisticRegression

lr = LogisticRegression(max_iter=200)
lr.fit(X_train, y_train)
print('Test accuracy:', lr.score(X_test, y_test))
Test accuracy: 1.0

Step 5: Clustering with KMeans#

Clustering groups baskets with similar purchase patterns.

This is useful for finding customer segments automatically.

Let us cluster a sample using the KMeans algorithm from scikit-learn.

We will cluster using a smaller set of item columns for speed.

# Pick top 10 frequent items for clustering
cols = all_items.value_counts().head(10).index.tolist()
sample_basket = basket_matrix[cols]
from sklearn.cluster import KMeans
kmeans = KMeans(n_clusters=3, random_state=42)
clusters = kmeans.fit_predict(sample_basket)
sample_basket['cluster'] = clusters
print(sample_basket.groupby('cluster').mean())
         whole milk  other vegetables  rolls/buns      soda    yogurt  \
cluster                                                                 
0          0.475446          1.000000         0.0  0.311384  0.295759   
1          0.405125          0.000000         0.0  0.290421  0.246492   
2          0.510638          0.419663         1.0  0.342627  0.318415   

         root vegetables  tropical fruit  bottled water   sausage  \
cluster                                                             
0               0.229911        0.231027       0.248884  0.222098   
1               0.206833        0.219646       0.183649  0.172666   
2               0.259721        0.252384       0.226706  0.235510   

         citrus fruit  
cluster                
0            0.194196  
1            0.164124  
2            0.205429  

Step 6: Mini ProjectMarket Basket Insights#

Practice: Create your own rule or cluster analysis.

Try changing the item or support threshold and see what new patterns appear.

You can analyze a different top item, or use more or fewer clusters in KMeans.

Exploring parameters is the best way to learn!

# Try it: predict another item's purchase
item = input('Enter an item name to predict (e.g., soda): ')
if item in basket_matrix.columns:
    from sklearn.linear_model import LogisticRegression
    y2 = basket_matrix[item]
    X2 = basket_matrix.drop([item, 'buys_whole_milk'], axis=1, errors='ignore')
    X2_train, X2_test, y2_train, y2_test = train_test_split(X2, y2, test_size=0.25, random_state=42)
    lr2 = LogisticRegression(max_iter=100)
    lr2.fit(X2_train, y2_train)
    print(f'Accuracy in predicting {item}:', lr2.score(X2_test, y2_test))
else:
    print('Item not found. Try a different name.')
    
Item not found. Try a different name.
# Explore: how many clusters work best?
for k in range(2, 6):
    kmeans_k = KMeans(n_clusters=k, random_state=42).fit(sample_basket[cols])
    print(f'k={k}: Inertia={kmeans_k.inertia_:.2f}')
    
k=2: Inertia=6732.80
k=3: Inertia=6175.64
k=4: Inertia=5925.07
k=5: Inertia=5722.58

Troubleshooting and Best Practices#

If something does not work, check for typos and missing dependencies.

Always inspect your data before trusting results.

Test your logic with small samples to spot errors early.

Document your process in comments and markdown cells.

Keep your code tidy and backed up.

Extra Tips#

Try more datasets, like movies or sports, for more practice.

Many data mining concepts also work for text or images.

Explore the scikit-learn and pandas documentation for more features.

Sharing your findings with friends helps you learn faster!

Challenge Exercises#

  1. Find a set of three items that often appear together.

  2. Predict a different item's purchase, such as 'soda' or 'yogurt.'

  3. Cluster with five or more clusters and describe what you observe.

  4. Try visualizing the basket sizes with a histogram.

Share your solutions in the comments to help other learners!

Recap#

In this lesson, you learned:

How to explore, clean, and prepare transactional data. How to group transactions into baskets for mining. Basic association rule mining for shopping pattern discovery. Simple classification and clustering from real data. Why testing, visualization, and parameter searching matter.

Keep practicing with new data to build your data mining confidence!

Thank you for learning Exploratory Data Mining!#

If you enjoyed this lesson, please like, subscribe, and share.

Check out the next video for deeper dives into machine learning!

Happy mining!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.