Lesson 13 · Data Mining
Understanding Histograms, Scatterplots, and Heatmaps for Data Visualization
In this lesson, you will learn how to visualize data using histograms, scatterplots, and heatmaps. We will use the Groceries Dataset to explore these plots.…
- CourseData Mining
- Lesson13 of 31
- Video16 min
- FormatJupyter notebook · 11 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbWeek 4: Data Visualization in Python#
In this lesson, you will learn how to visualize data using histograms, scatterplots, and heatmaps.
We will use the Groceries Dataset to explore these plots.
Visualizations help you see patterns and trends that are hard to spot in raw data.
Let us begin our journey to make data come alive!
import warnings; warnings.filterwarnings("ignore")
# Data setup
import pandas as pd
import kagglehub, os
import numpy as np
np.random.seed(42)
path = kagglehub.dataset_download('heeraldedhia/groceries-dataset')
files = os.listdir(path)
csv_file = [f for f in files if f.endswith('.csv')][0]
df = pd.read_csv(os.path.join(path, csv_file))
print(df.shape)
print(df.head(3))
Why Data Visualization Matters#
- Visualizations communicate information clearly.
- They help you identify patterns, spot outliers, and form new questions.
Think of visualizations like the map for your data journey.
# Let us see the basic structure of our dataset
df.info()
# Check for missing values
missing = df.isnull().sum()
print(missing)
# View top purchased items
top_items = df['itemDescription'].value_counts().head(10)
print(top_items)
# Visualize most popular items with a histogram
import matplotlib.pyplot as plt
plt.figure(figsize=(10,4))
top_items.plot(kind='bar', color='skyblue')
plt.ylabel('Frequency')
plt.title('Top 10 Purchased Grocery Items')
plt.show()
Understanding Histograms#
Histograms show how values are distributed.
- Bar charts work for categories, like item names.
- For numbers, histograms reveal the spread and shape.
Next, let us make a histogram from a numeric column.
# How many items does each customer buy per shopping trip?
items_per_trip = df.groupby(['Member_number', 'Date']).size()
plt.hist(items_per_trip, bins=20, color='purple', edgecolor='white')
plt.xlabel('Number of Items per Trip')
plt.ylabel('Number of Transactions')
plt.title('Distribution of Items Bought Each Trip')
plt.show()
Intro to Scatterplots#
Scatterplots show the relationship between two variables.
Each point is a pair of values for one example.
Let us look for patterns in a scatterplot!
# Scatterplot: Items per trip over time for one customer
sample_member = df['Member_number'].sample(1, random_state=42).iloc[0]
member_transactions = df[df['Member_number'] == sample_member]
trips = member_transactions.groupby('Date').size().reset_index(name='ItemsBought')
plt.scatter(trips['Date'], trips['ItemsBought'], alpha=0.7, color='orange')
plt.xlabel('Date')
plt.ylabel('Items Bought in Trip')
plt.title(f'Shopping Pattern Over Time: Member {sample_member}')
plt.xticks(rotation=45)
plt.show()
# Preprocess data for a heatmap: Days and items
item_by_day = df.groupby(['Date', 'itemDescription']).size().unstack(fill_value=0)
item_totals = item_by_day.sum(axis=0).sort_values(ascending=False)[:8]
small = item_by_day[item_totals.index]
small = small.T
# Draw the heatmap
import seaborn as sns
plt.figure(figsize=(12,5))
sns.heatmap(small, cmap='YlGnBu', linewidths=0.5, cbar_kws={'label': 'Count'})
plt.ylabel('Item')
plt.xlabel('Date')
plt.title('Heatmap of Item Purchases Across Days')
plt.show()
When to Use Each Visualization#
- Use histograms or bar charts for frequency distributions.
- Choose scatterplots for comparing two pairs of values.
- Pick a heatmap to spot clusters, patterns, or gaps in data across two categories (like time and item).
Try to choose the plot that matches the question you want to answer.
# Practice: Make your own histogram
your_col = input("Type column name: 'Member_number' or 'itemDescription': ")
vc = df[your_col].value_counts().head(10)
vc.plot(kind='bar', color='salmon')
plt.title(f'Your Top 10: {your_col}')
plt.show()
Best Practices#
- Label your axes and titles clearly.
- Do not overload plots with too much data at once.
- Pick color palettes that are easy to see.
Simple, clear plots are more powerful.
# Challenge: Try a heatmap for days and members
pivot = df.pivot_table(index='Date', columns='Member_number', values='itemDescription', aggfunc='count', fill_value=0)
sample_cols = pivot.columns[:15]
plt.figure(figsize=(12,8))
sns.heatmap(pivot[sample_cols], cmap='mako', cbar_kws={'label': 'Items Bought'})
plt.title('Shopping Heatmap: Dates vs. 15 Members')
plt.ylabel('Date')
plt.xlabel('Member Number')
plt.show()
Recap: Making Sense of the Data#
- Histograms and bar charts show what is common or rare.
- Scatterplots help reveal trends between two things.
- Heatmaps let you see many relationships at once.
You are now ready to tell your own data stories with pictures!
Next Steps#
- Practice by visualizing data from your own life or interests.
- Keep exploring! Try pie charts, boxplots, or line graphs.
If you enjoyed this lesson, like and subscribe for more beginner data lessons!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



