Mathew K Analytics

Lesson 12 · Data Mining

Understanding Summarization and Descriptive Statistics in Data Mining

Welcome! This week we explore how to make sense of your data using descriptive statistics. We will use the Groceries Dataset and learn key skills for…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Week 4: Summarization and Descriptive Statistics#

Welcome! This week we explore how to make sense of your data using descriptive statistics.

We will use the Groceries Dataset and learn key skills for summarizing real-world datasets.

By the end, you will know how to describe, visualize, and get basic insights from complex data.

# Suppress warning messages for a clean workspace
import warnings; warnings.filterwarnings("ignore")
import numpy as np
np.random.seed(42)
# Data setup (Groceries Dataset)
import pandas as pd
import kagglehub, os
path = kagglehub.dataset_download('heeraldedhia/groceries-dataset')
files = os.listdir(path)
csv_file = [f for f in files if f.endswith('.csv')][0]
df = pd.read_csv(os.path.join(path, csv_file))
print(df.shape)
print(df.head(3))
(38765, 3)
   Member_number        Date itemDescription
0           1808  21-07-2015  tropical fruit
1           2552  05-01-2015      whole milk
2           2300  19-09-2015       pip fruit

What does the Groceries Dataset look like?#

This dataset shows shoppers, dates, and the items they bought. Each row is a single item that a customer purchased on a specific date.

For data mining, we want to learn from these shopping patterns!

# Basic info: columns, types, and non-null values
df.info()
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 38765 entries, 0 to 38764
Data columns (total 3 columns):
 #   Column           Non-Null Count  Dtype 
---  ------           --------------  ----- 
 0   Member_number    38765 non-null  int64 
 1   Date             38765 non-null  object
 2   itemDescription  38765 non-null  object
dtypes: int64(1), object(2)
memory usage: 908.7+ KB
# Quick look at unique values for item and members
print("Unique items:", df['itemDescription'].nunique())
print("Unique customers:", df['Member_number'].nunique())
Unique items: 167
Unique customers: 3898
# Are there missing values?
print(df.isnull().sum())
Member_number      0
Date               0
itemDescription    0
dtype: int64
# How many items does each customer buy per trip?
basket_sizes = df.groupby(['Member_number']).size()
print(basket_sizes.describe())
count    3898.000000
mean        9.944844
std         5.310796
min         2.000000
25%         6.000000
50%         9.000000
75%        13.000000
max        36.000000
dtype: float64
# What is the most popular item?
top_items = df['itemDescription'].value_counts().head(5)
print(top_items)
itemDescription
whole milk          2502
other vegetables    1898
rolls/buns          1716
soda                1514
yogurt              1334
Name: count, dtype: int64
# How many total transactions (shopping trips) are there?
num_transactions = df.groupby(['Member_number', 'Date']).ngroups
print("Total shopping trips:", num_transactions)
Total shopping trips: 14963

Exploring: Summarizing with Descriptive Statistics#

Descriptive statistics help us quickly understand the "shape" of our data.

Common summary stats include:

  • Count: how many
  • Mean: the average
  • Median: the middle value
  • Standard deviation: how spread out the numbers are
  • Minimum and maximum: the smallest and largest values

Let us begin with these basics.

# Summary stats for basket sizes
print("Count, Mean, Std, Min, 25%, 50%, 75%, Max:")
print(basket_sizes.describe())
Count, Mean, Std, Min, 25%, 50%, 75%, Max:
count    3898.000000
mean        9.944844
std         5.310796
min         2.000000
25%         6.000000
50%         9.000000
75%        13.000000
max        36.000000
dtype: float64
# What day do people shop the most?
busy_days = df['Date'].value_counts().sort_values(ascending=False).head(3)
print(busy_days)
Date
21-01-2015    96
21-07-2015    93
08-08-2015    92
Name: count, dtype: int64
# Visualize basket size distribution
import matplotlib.pyplot as plt
plt.figure(figsize=(6,3))
basket_sizes.plot(kind='hist', bins=range(1,16), color='skyblue', edgecolor='black')
plt.title('Number of Items per Shopping Trip')
plt.xlabel('Number of Items')
plt.ylabel('Number of Trips')
plt.tight_layout()
plt.show()
No description has been provided for this image
# Find the median number of items bought per customer overall
median_basket = basket_sizes.groupby('Member_number').median()
print(median_basket.describe())
count    3898.000000
mean        9.944844
std         5.310796
min         2.000000
25%         6.000000
50%         9.000000
75%        13.000000
max        36.000000
dtype: float64
# Top 3 items for a specific customer
customer_id = int(input("Enter a customer number, e.g. 1808: "))
top_for_customer = df[df['Member_number'] == customer_id]['itemDescription'].value_counts().head(3)
print(top_for_customer)
itemDescription
tropical fruit              1
long life bakery product    1
meat                        1
Name: count, dtype: int64
# Build a table: each customer's average basket size
avg_basket_size = basket_sizes.groupby('Member_number').mean()
avg_basket_size = avg_basket_size.rename('Average Basket Size')
print(avg_basket_size.head())
Member_number
1000    13.0
1001    12.0
1002     8.0
1003     8.0
1004    21.0
Name: Average Basket Size, dtype: float64
# Box plot: variation in customer basket sizes
plt.figure(figsize=(5,3))
avg_basket_size.plot(kind='box')
plt.ylabel('Average Items per Basket')
plt.title('Basket Size Spread across Customers')
plt.tight_layout()
plt.show()
No description has been provided for this image
# Median, min, and max basket size by customer
basket_stats = basket_sizes.groupby('Member_number').agg(['median', 'min', 'max'])
print(basket_stats.head())
               median  min  max
Member_number                  
1000             13.0   13   13
1001             12.0   12   12
1002              8.0    8    8
1003              8.0    8    8
1004             21.0   21   21
# Which customer buys the largest baskets?
largest_basket_customers = avg_basket_size.sort_values(ascending=False).head(3)
print(largest_basket_customers)
Member_number
3180    36.0
3050    33.0
2051    33.0
Name: Average Basket Size, dtype: float64

Wrap-up and What Next?#

You have learned to use descriptive statistics and summaries to learn from data.

You practiced with the Groceries dataset and saw how simple stats can reveal powerful insights.

Try more questions! Can you:

  • Find customers who rarely visit?
  • Discover which items are bought together?
  • Spot shopping habits during holidays?

Next week: Association rules and market basket analysis!

Like and subscribe for more data mining tutorials!

Challenge: Practice and Share!#

  • Try answering: Who are the most loyal customers?
  • Can you make a new chart or summary with matplotlib or pandas?

Share your code in the comments. We love to see your progress!

See you in the next lesson!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.