Mathew K Analytics

Lesson 26 · Data Mining

Understanding Density-Based Clustering with DBSCAN: Principles and Python Implementation

Welcome to your introduction to DBSCAN clustering! You will learn: What is density-based clustering. Why DBSCAN is different from K-Means. How to find…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Week 7: Density-Based Clustering with DBSCAN#

Welcome to your introduction to DBSCAN clustering!

You will learn:

  • What is density-based clustering.
  • Why DBSCAN is different from K-Means.
  • How to find groups in real-world data even if groups are not circles!

We will use the Mall Customers Dataset to practice.

Let us dive in!

# Data setup (Mall Customers Dataset)
import warnings; warnings.filterwarnings("ignore")
import pandas as pd
import numpy as np
np.random.seed(42)
url = 'https://gist.githubusercontent.com/pravalliyaram/5c05f43d2351249927b8a3f3cc3e5ecf/raw/Mall_Customers.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
(200, 5)
   CustomerID  Gender  Age  Annual Income (k$)  Spending Score (1-100)
0           1    Male   19                  15                      39
1           2    Male   21                  15                      81
2           3  Female   20                  16                       6

What is clustering?#

Clustering is finding groups (clusters) of similar data points.

Why do we do this?

  • To discover patterns for example, what types of customers shop here?
  • To divide customers for marketing, product planning, or help guides.

Clustering is a basic way to explore structure in data with no labels needed.

# Preview relevant columns
print(df.columns)
print(df[['Age','Annual Income (k$)','Spending Score (1-100)']].describe())
Index(['CustomerID', 'Gender', 'Age', 'Annual Income (k$)',
       'Spending Score (1-100)'],
      dtype='object')
              Age  Annual Income (k$)  Spending Score (1-100)
count  200.000000          200.000000              200.000000
mean    38.850000           60.560000               50.200000
std     13.969007           26.264721               25.823522
min     18.000000           15.000000                1.000000
25%     28.750000           41.500000               34.750000
50%     36.000000           61.500000               50.000000
75%     49.000000           78.000000               73.000000
max     70.000000          137.000000               99.000000

K-Means vs. DBSCAN#

You may have heard of K-Means.

K-Means forms round, equally-sized clusters.

But what if groups are messy? DBSCAN to the rescue!

  • DBSCAN stands for Density-Based Spatial Clustering of Applications with Noise.
  • It finds clusters of any shape, even with outliers.
  • You do not need to pick the number of clusters first.
# Visualize customers by incomes and spending
import matplotlib.pyplot as plt
plt.scatter(df['Annual Income (k$)'], df['Spending Score (1-100)'], alpha=0.6)
plt.xlabel('Annual Income (k$)')
plt.ylabel('Spending Score (1-100)')
plt.title('Mall Customers: Income vs. Spending')
plt.show()
No description has been provided for this image

What is density?#

In DBSCAN, 'density' means areas where lots of points are close together.

  • A cluster has many points in each others neighborhoods.
  • Points with few nearby neighbors may be called 'noise' (outliers).

DBSCAN needs two settings:

  • eps: how close points must be to consider as neighbors.
  • min_samples: how many points needed to make a cluster.
# Prepare data for clustering
X = df[['Annual Income (k$)', 'Spending Score (1-100)']].copy()
# Standardize features
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)

Fitting DBSCAN#

We will use sklearns DBSCAN to find clusters.

Let us start with some basic values for eps and min_samples.

Experimentation is normal! DBSCAN gives different results if you tweak eps.

# Run DBSCAN clustering
from sklearn.cluster import DBSCAN
dbscan = DBSCAN(eps=0.5, min_samples=5)
labels = dbscan.fit_predict(X_scaled)
print(pd.Series(labels).value_counts())
 0    157
 1     35
-1      8
Name: count, dtype: int64
# Plot DBSCAN results
plt.figure(figsize=(7,5))
colors = ['red','green','blue','purple','orange','pink','gray']
for k in set(labels):
    color = 'black' if k == -1 else colors[k%len(colors)]
    class_member_mask = (labels == k)
    plt.scatter(X[class_member_mask]['Annual Income (k$)'], 
                X[class_member_mask]['Spending Score (1-100)'],
                c=color, label='Cluster '+str(k) if k!=-1 else 'Noise', alpha=0.5)
plt.xlabel('Annual Income (k$)')
plt.ylabel('Spending Score (1-100)')
plt.title('DBSCAN Clusters')
plt.legend()
plt.show()
No description has been provided for this image
# Change parameters to explore more
eps = float(input("Try a new eps (example 0.2 to 1.0): "))
min_samples = int(input("Try a new min_samples (example 3 to 10): "))
dbscan2 = DBSCAN(eps=eps, min_samples=min_samples)
labels2 = dbscan2.fit_predict(X_scaled)
plt.figure(figsize=(7,5))
for k in set(labels2):
    color = 'black' if k == -1 else colors[k%len(colors)]
    class_member_mask = (labels2 == k)
    plt.scatter(X[class_member_mask]['Annual Income (k$)'], 
                X[class_member_mask]['Spending Score (1-100)'],
                c=color, label='Cluster '+str(k) if k!=-1 else 'Noise', alpha=0.5)
plt.xlabel('Annual Income (k$)')
plt.ylabel('Spending Score (1-100)')
plt.title('DBSCAN: Your Settings')
plt.legend()
plt.show()
No description has been provided for this image

Understanding DBSCAN Labels#

Each point gets a label like 0, 1, 2 ... or -1.

-1 means the point is 'noise' it is not close to any cluster.

Inspecting and counting cluster sizes helps you tell if your clustering worked.

# Get a summary of each group
import numpy as np
for k in set(labels):
    class_mask = (labels == k)
    print(f"Cluster {k if k!=-1 else 'Noise'}: Count: {class_mask.sum()}")
    print(' Mean income:', np.round(X[class_mask]['Annual Income (k$)'].mean(),2),
          'Mean score:', np.round(X[class_mask]['Spending Score (1-100)'].mean(),2))
    print('-----')
    
Cluster 0: Count: 157
 Mean income: 52.49 Mean score: 43.1
-----
Cluster 1: Count: 35
 Mean income: 82.54 Mean score: 82.8
-----
Cluster Noise: Count: 8
 Mean income: 122.75 Mean score: 46.88
-----
# Comparing DBSCAN to KMeans
from sklearn.cluster import KMeans
kmeans = KMeans(n_clusters=3, random_state=42)
k_labels = kmeans.fit_predict(X_scaled)
plt.figure(figsize=(7,5))
for k in set(k_labels):
    class_mask = (k_labels == k)
    plt.scatter(X[class_mask]['Annual Income (k$)'], X[class_mask]['Spending Score (1-100)'],
                c=colors[k%len(colors)], label='KMeans '+str(k), alpha=0.5)
plt.xlabel('Annual Income (k$)')
plt.ylabel('Spending Score (1-100)')
plt.title('K-Means Clustering (for comparison)')
plt.legend()
plt.show()
No description has been provided for this image

Best practices for DBSCAN#

  • Always scale your features first.
  • Try several eps and min_samples values to get good results.
  • Look at both the cluster plot and group summaries.
  • Watch out for all-noise or too-large clusters adjust settings if you see these.

Remember: DBSCAN works best when clusters have dense, separate groups.

# Troubleshooting DBSCAN clustering
if (labels == -1).all():
    print("All data points are marked as noise. Try increasing eps.")
elif len(set(labels)) == 1 and -1 not in set(labels):
    print("Only one big cluster found. Try lowering eps or increasing min_samples.")
else:
    print("DBSCAN found:", len(set(labels)) - (1 if -1 in set(labels) else 0), "clusters and", sum(labels==-1), "noise points.")
    
DBSCAN found: 2 clusters and 8 noise points.
# Extra: Automatic parameter search using NearestNeighbors
from sklearn.neighbors import NearestNeighbors
neighbors = NearestNeighbors(n_neighbors=5)
neighbors_fit = neighbors.fit(X_scaled)
distances, indices = neighbors_fit.kneighbors(X_scaled)
distances = np.sort(distances[:,4])
plt.plot(distances)
plt.title('K-distance Graph (for eps guess)')
plt.xlabel('Points sorted by distance')
plt.ylabel('5th Nearest Neighbor distance')
plt.show()
No description has been provided for this image

Mini-Challenge: Cluster customer types#

Can you find parameter values for DBSCAN that give at least two clusters and less than 20% noise?

  • Try different eps values.
  • Try different min_samples.
  • Look at the summary printout and plots.

This is good practice for when you try DBSCAN on your own data!

Recap#

Today you learned:

  • What DBSCAN is and why it is different than KMeans.
  • How to try, critique, and tune its clustering.
  • That scaling data is important!

Keep experimenting data mining is a skill you build with practice.

Thanks for learning!#

If you enjoyed this lesson, please like, subscribe, and share.

Try using DBSCAN on your own dataset and let us know your results in the comments below!

Happy clustering see you next time!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.