Mathew K Analytics

Lesson 23 · Python For Machine Learning

Master DBSCAN & Hierarchical Clustering in Python for Unsupervised Learning

Clustering is an important technique for finding groups in unlabeled data. This lesson will help you understand and use two powerful clustering methods:…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb
 

Welcome to DBSCAN and Hierarchical Clustering in Python!#

Clustering is an important technique for finding groups in unlabeled data.

This lesson will help you understand and use two powerful clustering methods:

  • DBSCAN: A density-based approach.
  • Hierarchical clustering: Builds a tree based on data similarity.

We will use real-world data and friendly explanations.

Let us get started!

# First, let us import what we need and silence warnings
import warnings
warnings.filterwarnings("ignore")
import numpy as np
import pandas as pd

What is Clustering?#

Clustering tries to group similar data points together.

It is useful for:

  • Customer segmentation in business.
  • Analyzing images or text.
  • Discovering patterns in health or science data.

The main point: Data points inside the same group (cluster) should be similar!

# Data setup
# Let us use the Mall Customers dataset: shopping data for clustering.
url = "https://raw.githubusercontent.com/rithwik00/Mall-Customer-Dataset/main/Mall_Customers.csv"
df = pd.read_csv(url)
print("Shape:", df.shape)
df.head()
Shape: (200, 5)
CustomerID Gender Age Annual Income (k$) Spending Score (1-100)
0 1 Male 19 15 39
1 2 Male 21 15 81
2 3 Female 20 16 6
3 4 Female 23 16 77
4 5 Female 31 17 40
# Let us look at the column names and some summary statistics too.
df.columns = df.columns.str.strip()
print("Columns:", df.columns.tolist())
df.describe()
Columns: ['CustomerID', 'Gender', 'Age', 'Annual Income (k$)', 'Spending Score (1-100)']
CustomerID Age Annual Income (k$) Spending Score (1-100)
count 200.000000 200.000000 200.000000 200.000000
mean 100.500000 38.850000 60.560000 50.200000
std 57.879185 13.969007 26.264721 25.823522
min 1.000000 18.000000 15.000000 1.000000
25% 50.750000 28.750000 41.500000 34.750000
50% 100.500000 36.000000 61.500000 50.000000
75% 150.250000 49.000000 78.000000 73.000000
max 200.000000 70.000000 137.000000 99.000000
# We will focus on two columns: 'Annual Income (k$)' and 'Spending Score (1-100)'
features = df[["Annual Income (k$)", "Spending Score (1-100)"]].copy()
features.head()
Annual Income (k$) Spending Score (1-100)
0 15 39
1 15 81
2 16 6
3 16 77
4 17 40

What is DBSCAN?#

DBSCAN is a clustering method based on how closely packed points are.

  • Groups points that are near each other into clusters.
  • Can find clusters of any shape.
  • Marks points that are too far away as 'noise'.

We do not need to tell DBSCAN how many clusters to find.

# Let us try DBSCAN
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import DBSCAN

# First, scale our features so each has equal weight
scaler = StandardScaler()
X_scaled = scaler.fit_transform(features)

# Set up and fit DBSCAN
dbscan = DBSCAN(eps=0.5, min_samples=5)
labels = dbscan.fit_predict(X_scaled)
# Add cluster labels to our data
df["Cluster"] = labels
df[["Annual Income (k$)", "Spending Score (1-100)", "Cluster"]].head()
Annual Income (k$) Spending Score (1-100) Cluster
0 15 39 0
1 15 81 0
2 16 6 0
3 16 77 0
4 17 40 0

Checking the Results#

Let us see how many clusters DBSCAN found.

DBSCAN uses -1 to mark 'noise' points that do not fit in any group.

# Count how many points are in each cluster
print(df["Cluster"].value_counts())
Cluster
 0    157
 1     35
-1      8
Name: count, dtype: int64
# Let us visualize our DBSCAN clusters
import matplotlib.pyplot as plt

plt.figure(figsize=(8,6))
for cluster in sorted(df["Cluster"].unique()):
    mask = df["Cluster"] == cluster
    if cluster == -1:
        color = "gray"
        label = "Noise"
    else:
        color = None
        label = f"Cluster {cluster}"
    plt.scatter(df.loc[mask, "Annual Income (k$)"],
                df.loc[mask, "Spending Score (1-100)"],
                c=color, label=label, alpha=0.6)
plt.xlabel("Annual Income (k$)")
plt.ylabel("Spending Score (1-100)")
plt.title("DBSCAN clustering results")
plt.legend()
plt.show()
No description has been provided for this image

What is Hierarchical Clustering?#

This method groups points by building a tree that shows how close data points are.

  • Agglomerative: Start with every point separate, then merge the closest pair again and again.
  • Results can be shown in a tree-like chart called a dendrogram.

You can pick the number of clusters by 'cutting' the tree.

# Let us try agglomerative hierarchical clustering
from sklearn.cluster import AgglomerativeClustering
clustering = AgglomerativeClustering(n_clusters=3)
hier_labels = clustering.fit_predict(X_scaled)
df["HierCluster"] = hier_labels
df[["Annual Income (k$)", "Spending Score (1-100)", "HierCluster"]].head()
Annual Income (k$) Spending Score (1-100) HierCluster
0 15 39 0
1 15 81 0
2 16 6 0
3 16 77 0
4 17 40 0
# Visualize hierarchical clustering results
plt.figure(figsize=(8,6))
for cluster in sorted(df["HierCluster"].unique()):
    mask = df["HierCluster"] == cluster
    plt.scatter(df.loc[mask, "Annual Income (k$)"],
                df.loc[mask, "Spending Score (1-100)"],
                label=f"Cluster {cluster}", alpha=0.7)
plt.xlabel("Annual Income (k$)")
plt.ylabel("Spending Score (1-100)")
plt.title("Hierarchical Clustering Results")
plt.legend()
plt.show()
No description has been provided for this image
# Draw a dendrogram to see how points were grouped step by step
from scipy.cluster.hierarchy import dendrogram, linkage
linked = linkage(X_scaled, method="ward")
plt.figure(figsize=(12,6))
dendrogram(linked, orientation="top", distance_sort="descending", show_leaf_counts=False)
plt.title("Dendrogram - Hierarchical Clustering")
plt.xlabel("Sample index")
plt.ylabel("Distance")
plt.show()
No description has been provided for this image

DBSCAN vs Hierarchical Clustering#

Both algorithms help us find groups, but they have different strengths.

  • DBSCAN is great for data that has weird shapes and noise.
  • Hierarchical clustering gives us a tree view to pick our number of clusters.

Try both to see what works best for your data.

# Handling input: Let us ask the user to choose a method
print("Which clustering method would you like to try? Type 'dbscan' or 'hierarchical':")
method = input()
if method == 'dbscan':
    result = df["Cluster"]
    print("DBSCAN clusters selected.")
elif method == 'hierarchical':
    result = df["HierCluster"]
    print("Hierarchical clusters selected.")
else:
    print("Unknown option. Please run again and type 'dbscan' or 'hierarchical'.")
    
Which clustering method would you like to try? Type 'dbscan' or 'hierarchical':
DBSCAN clusters selected.
 
# Let us look at which labels the user got
if method == 'dbscan':
    print(result.value_counts())
elif method == 'hierarchical':
    print(result.value_counts())

    
Cluster
 0    157
 1     35
-1      8
Name: count, dtype: int64

Mini Project: Segmenting Customers#

Let us pretend we work for the mall.

Our task is to use clustering to find distinct groups of shoppers.

We want to discover customer types for promotions.

Can you describe one customer group you find?

# Print statistics for each cluster
choice = 0 if method == 'dbscan' else 1
clusters = ["Cluster", "HierCluster"]
for label in sorted(df[clusters[choice]].unique()):
    print(f"Stats for {clusters[choice]} {label}: ")
    print(df[df[clusters[choice]] == label][["Age", "Annual Income (k$)", "Spending Score (1-100)"]].describe())
    print()

    
Stats for Cluster -1: 
             Age  Annual Income (k$)  Spending Score (1-100)
count   8.000000            8.000000                8.000000
mean   35.750000          122.750000               46.875000
std     6.497252           11.510864               32.108911
min    30.000000          103.000000                8.000000
25%    32.000000          118.250000               17.500000
50%    32.500000          123.000000               48.500000
75%    37.500000          128.750000               75.250000
max    47.000000          137.000000               83.000000

Stats for Cluster 0: 
              Age  Annual Income (k$)  Spending Score (1-100)
count  157.000000          157.000000              157.000000
mean    40.369427           52.490446               43.101911
std     15.249332           21.811141               22.249225
min     18.000000           15.000000                1.000000
25%     26.000000           37.000000               28.000000
50%     40.000000           54.000000               46.000000
75%     51.000000           65.000000               55.000000
max     70.000000          103.000000               99.000000

Stats for Cluster 1: 
             Age  Annual Income (k$)  Spending Score (1-100)
count  35.000000           35.000000               35.000000
mean   32.742857           82.542857               82.800000
std     3.890735           10.925800                9.498607
min    27.000000           69.000000               63.000000
25%    30.000000           74.500000               75.000000
50%    32.000000           78.000000               86.000000
75%    36.000000           87.500000               90.500000
max    40.000000          113.000000               97.000000

# Extra tip: Remove noise or outliers from DBSCAN if needed
if "Cluster" in df and -1 in df["Cluster"].values:
    clean_df = df[df["Cluster"] != -1].copy()
    print(f"Rows after removing noise: {clean_df.shape[0]}")
else:
    clean_df = df.copy()

    
Rows after removing noise: 192
# Challenge: Try DBSCAN with different eps values using input()
eps = float(input("Type a value for eps (try 0.2, 0.4, 0.8): "))
db_alt = DBSCAN(eps=eps, min_samples=5)
labels_alt = db_alt.fit_predict(X_scaled)
print("Clusters with eps =", eps, ":", np.unique(labels_alt))
Clusters with eps = 0.8 : [0]
 
# Troubleshooting: What if you get too many or too few clusters?
# - Lower eps means fewer neighbors needed, so it makes more clusters.
# - Raise eps if you get too much noise.
# - Lower min_samples to allow smaller groups.
# Use .fit_predict again to update clusters.

Recap: What did we learn?#

  • DBSCAN groups nearby points and finds noise automatically.
  • Hierarchical clustering builds a tree of groupings.
  • Visualizing results helps you check if clusters match what you expect.

Clustering can reveal hidden structures and simplify data!

Keep Learning More!#

Try clustering with your own data, or experiment with more features.

Check out more Python and machine learning tutorials on our YouTube channel!

Leave a comment with what kind of data you would like to cluster next.

Happy coding!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.