Lesson 12 · Data Mining
Understanding Summarization and Descriptive Statistics in Data Mining
Welcome! This week we explore how to make sense of your data using descriptive statistics. We will use the Groceries Dataset and learn key skills for…
- CourseData Mining
- Lesson12 of 31
- Video16 min
- FormatJupyter notebook · 17 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbWeek 4: Summarization and Descriptive Statistics#
Welcome! This week we explore how to make sense of your data using descriptive statistics.
We will use the Groceries Dataset and learn key skills for summarizing real-world datasets.
By the end, you will know how to describe, visualize, and get basic insights from complex data.
# Suppress warning messages for a clean workspace
import warnings; warnings.filterwarnings("ignore")
import numpy as np
np.random.seed(42)
# Data setup (Groceries Dataset)
import pandas as pd
import kagglehub, os
path = kagglehub.dataset_download('heeraldedhia/groceries-dataset')
files = os.listdir(path)
csv_file = [f for f in files if f.endswith('.csv')][0]
df = pd.read_csv(os.path.join(path, csv_file))
print(df.shape)
print(df.head(3))
What does the Groceries Dataset look like?#
This dataset shows shoppers, dates, and the items they bought. Each row is a single item that a customer purchased on a specific date.
For data mining, we want to learn from these shopping patterns!
# Basic info: columns, types, and non-null values
df.info()
# Quick look at unique values for item and members
print("Unique items:", df['itemDescription'].nunique())
print("Unique customers:", df['Member_number'].nunique())
# Are there missing values?
print(df.isnull().sum())
# How many items does each customer buy per trip?
basket_sizes = df.groupby(['Member_number']).size()
print(basket_sizes.describe())
# What is the most popular item?
top_items = df['itemDescription'].value_counts().head(5)
print(top_items)
# How many total transactions (shopping trips) are there?
num_transactions = df.groupby(['Member_number', 'Date']).ngroups
print("Total shopping trips:", num_transactions)
Exploring: Summarizing with Descriptive Statistics#
Descriptive statistics help us quickly understand the "shape" of our data.
Common summary stats include:
- Count: how many
- Mean: the average
- Median: the middle value
- Standard deviation: how spread out the numbers are
- Minimum and maximum: the smallest and largest values
Let us begin with these basics.
# Summary stats for basket sizes
print("Count, Mean, Std, Min, 25%, 50%, 75%, Max:")
print(basket_sizes.describe())
# What day do people shop the most?
busy_days = df['Date'].value_counts().sort_values(ascending=False).head(3)
print(busy_days)
# Visualize basket size distribution
import matplotlib.pyplot as plt
plt.figure(figsize=(6,3))
basket_sizes.plot(kind='hist', bins=range(1,16), color='skyblue', edgecolor='black')
plt.title('Number of Items per Shopping Trip')
plt.xlabel('Number of Items')
plt.ylabel('Number of Trips')
plt.tight_layout()
plt.show()
# Find the median number of items bought per customer overall
median_basket = basket_sizes.groupby('Member_number').median()
print(median_basket.describe())
# Top 3 items for a specific customer
customer_id = int(input("Enter a customer number, e.g. 1808: "))
top_for_customer = df[df['Member_number'] == customer_id]['itemDescription'].value_counts().head(3)
print(top_for_customer)
# Build a table: each customer's average basket size
avg_basket_size = basket_sizes.groupby('Member_number').mean()
avg_basket_size = avg_basket_size.rename('Average Basket Size')
print(avg_basket_size.head())
# Box plot: variation in customer basket sizes
plt.figure(figsize=(5,3))
avg_basket_size.plot(kind='box')
plt.ylabel('Average Items per Basket')
plt.title('Basket Size Spread across Customers')
plt.tight_layout()
plt.show()
# Median, min, and max basket size by customer
basket_stats = basket_sizes.groupby('Member_number').agg(['median', 'min', 'max'])
print(basket_stats.head())
# Which customer buys the largest baskets?
largest_basket_customers = avg_basket_size.sort_values(ascending=False).head(3)
print(largest_basket_customers)
Wrap-up and What Next?#
You have learned to use descriptive statistics and summaries to learn from data.
You practiced with the Groceries dataset and saw how simple stats can reveal powerful insights.
Try more questions! Can you:
- Find customers who rarely visit?
- Discover which items are bought together?
- Spot shopping habits during holidays?
Next week: Association rules and market basket analysis!
Like and subscribe for more data mining tutorials!
Challenge: Practice and Share!#
- Try answering: Who are the most loyal customers?
- Can you make a new chart or summary with matplotlib or pandas?
Share your code in the comments. We love to see your progress!
See you in the next lesson!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



