Lesson 23 · Python For Machine Learning
Master DBSCAN & Hierarchical Clustering in Python for Unsupervised Learning
Clustering is an important technique for finding groups in unlabeled data. This lesson will help you understand and use two powerful clustering methods:…
- CoursePython For Machine Learning
- Lesson23 of 16
- Video12 min
- FormatJupyter notebook · 19 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynb
Welcome to DBSCAN and Hierarchical Clustering in Python!#
Clustering is an important technique for finding groups in unlabeled data.
This lesson will help you understand and use two powerful clustering methods:
- DBSCAN: A density-based approach.
- Hierarchical clustering: Builds a tree based on data similarity.
We will use real-world data and friendly explanations.
Let us get started!
# First, let us import what we need and silence warnings
import warnings
warnings.filterwarnings("ignore")
import numpy as np
import pandas as pd
What is Clustering?#
Clustering tries to group similar data points together.
It is useful for:
- Customer segmentation in business.
- Analyzing images or text.
- Discovering patterns in health or science data.
The main point: Data points inside the same group (cluster) should be similar!
# Data setup
# Let us use the Mall Customers dataset: shopping data for clustering.
url = "https://raw.githubusercontent.com/rithwik00/Mall-Customer-Dataset/main/Mall_Customers.csv"
df = pd.read_csv(url)
print("Shape:", df.shape)
df.head()
# Let us look at the column names and some summary statistics too.
df.columns = df.columns.str.strip()
print("Columns:", df.columns.tolist())
df.describe()
# We will focus on two columns: 'Annual Income (k$)' and 'Spending Score (1-100)'
features = df[["Annual Income (k$)", "Spending Score (1-100)"]].copy()
features.head()
What is DBSCAN?#
DBSCAN is a clustering method based on how closely packed points are.
- Groups points that are near each other into clusters.
- Can find clusters of any shape.
- Marks points that are too far away as 'noise'.
We do not need to tell DBSCAN how many clusters to find.
# Let us try DBSCAN
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import DBSCAN
# First, scale our features so each has equal weight
scaler = StandardScaler()
X_scaled = scaler.fit_transform(features)
# Set up and fit DBSCAN
dbscan = DBSCAN(eps=0.5, min_samples=5)
labels = dbscan.fit_predict(X_scaled)
# Add cluster labels to our data
df["Cluster"] = labels
df[["Annual Income (k$)", "Spending Score (1-100)", "Cluster"]].head()
Checking the Results#
Let us see how many clusters DBSCAN found.
DBSCAN uses -1 to mark 'noise' points that do not fit in any group.
# Count how many points are in each cluster
print(df["Cluster"].value_counts())
# Let us visualize our DBSCAN clusters
import matplotlib.pyplot as plt
plt.figure(figsize=(8,6))
for cluster in sorted(df["Cluster"].unique()):
mask = df["Cluster"] == cluster
if cluster == -1:
color = "gray"
label = "Noise"
else:
color = None
label = f"Cluster {cluster}"
plt.scatter(df.loc[mask, "Annual Income (k$)"],
df.loc[mask, "Spending Score (1-100)"],
c=color, label=label, alpha=0.6)
plt.xlabel("Annual Income (k$)")
plt.ylabel("Spending Score (1-100)")
plt.title("DBSCAN clustering results")
plt.legend()
plt.show()
What is Hierarchical Clustering?#
This method groups points by building a tree that shows how close data points are.
- Agglomerative: Start with every point separate, then merge the closest pair again and again.
- Results can be shown in a tree-like chart called a dendrogram.
You can pick the number of clusters by 'cutting' the tree.
# Let us try agglomerative hierarchical clustering
from sklearn.cluster import AgglomerativeClustering
clustering = AgglomerativeClustering(n_clusters=3)
hier_labels = clustering.fit_predict(X_scaled)
df["HierCluster"] = hier_labels
df[["Annual Income (k$)", "Spending Score (1-100)", "HierCluster"]].head()
# Visualize hierarchical clustering results
plt.figure(figsize=(8,6))
for cluster in sorted(df["HierCluster"].unique()):
mask = df["HierCluster"] == cluster
plt.scatter(df.loc[mask, "Annual Income (k$)"],
df.loc[mask, "Spending Score (1-100)"],
label=f"Cluster {cluster}", alpha=0.7)
plt.xlabel("Annual Income (k$)")
plt.ylabel("Spending Score (1-100)")
plt.title("Hierarchical Clustering Results")
plt.legend()
plt.show()
# Draw a dendrogram to see how points were grouped step by step
from scipy.cluster.hierarchy import dendrogram, linkage
linked = linkage(X_scaled, method="ward")
plt.figure(figsize=(12,6))
dendrogram(linked, orientation="top", distance_sort="descending", show_leaf_counts=False)
plt.title("Dendrogram - Hierarchical Clustering")
plt.xlabel("Sample index")
plt.ylabel("Distance")
plt.show()
DBSCAN vs Hierarchical Clustering#
Both algorithms help us find groups, but they have different strengths.
- DBSCAN is great for data that has weird shapes and noise.
- Hierarchical clustering gives us a tree view to pick our number of clusters.
Try both to see what works best for your data.
# Handling input: Let us ask the user to choose a method
print("Which clustering method would you like to try? Type 'dbscan' or 'hierarchical':")
method = input()
if method == 'dbscan':
result = df["Cluster"]
print("DBSCAN clusters selected.")
elif method == 'hierarchical':
result = df["HierCluster"]
print("Hierarchical clusters selected.")
else:
print("Unknown option. Please run again and type 'dbscan' or 'hierarchical'.")
# Let us look at which labels the user got
if method == 'dbscan':
print(result.value_counts())
elif method == 'hierarchical':
print(result.value_counts())
Mini Project: Segmenting Customers#
Let us pretend we work for the mall.
Our task is to use clustering to find distinct groups of shoppers.
We want to discover customer types for promotions.
Can you describe one customer group you find?
# Print statistics for each cluster
choice = 0 if method == 'dbscan' else 1
clusters = ["Cluster", "HierCluster"]
for label in sorted(df[clusters[choice]].unique()):
print(f"Stats for {clusters[choice]} {label}: ")
print(df[df[clusters[choice]] == label][["Age", "Annual Income (k$)", "Spending Score (1-100)"]].describe())
print()
# Extra tip: Remove noise or outliers from DBSCAN if needed
if "Cluster" in df and -1 in df["Cluster"].values:
clean_df = df[df["Cluster"] != -1].copy()
print(f"Rows after removing noise: {clean_df.shape[0]}")
else:
clean_df = df.copy()
# Challenge: Try DBSCAN with different eps values using input()
eps = float(input("Type a value for eps (try 0.2, 0.4, 0.8): "))
db_alt = DBSCAN(eps=eps, min_samples=5)
labels_alt = db_alt.fit_predict(X_scaled)
print("Clusters with eps =", eps, ":", np.unique(labels_alt))
# Troubleshooting: What if you get too many or too few clusters?
# - Lower eps means fewer neighbors needed, so it makes more clusters.
# - Raise eps if you get too much noise.
# - Lower min_samples to allow smaller groups.
# Use .fit_predict again to update clusters.
Recap: What did we learn?#
- DBSCAN groups nearby points and finds noise automatically.
- Hierarchical clustering builds a tree of groupings.
- Visualizing results helps you check if clusters match what you expect.
Clustering can reveal hidden structures and simplify data!
Keep Learning More!#
Try clustering with your own data, or experiment with more features.
Check out more Python and machine learning tutorials on our YouTube channel!
Leave a comment with what kind of data you would like to cluster next.
Happy coding!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



