Lesson 26 · Data Mining
Understanding Density-Based Clustering with DBSCAN: Principles and Python Implementation
Welcome to your introduction to DBSCAN clustering! You will learn: What is density-based clustering. Why DBSCAN is different from K-Means. How to find…
- CourseData Mining
- Lesson26 of 31
- Video21 min
- FormatJupyter notebook · 12 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbWeek 7: Density-Based Clustering with DBSCAN#
Welcome to your introduction to DBSCAN clustering!
You will learn:
- What is density-based clustering.
- Why DBSCAN is different from K-Means.
- How to find groups in real-world data even if groups are not circles!
We will use the Mall Customers Dataset to practice.
Let us dive in!
# Data setup (Mall Customers Dataset)
import warnings; warnings.filterwarnings("ignore")
import pandas as pd
import numpy as np
np.random.seed(42)
url = 'https://gist.githubusercontent.com/pravalliyaram/5c05f43d2351249927b8a3f3cc3e5ecf/raw/Mall_Customers.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
What is clustering?#
Clustering is finding groups (clusters) of similar data points.
Why do we do this?
- To discover patterns for example, what types of customers shop here?
- To divide customers for marketing, product planning, or help guides.
Clustering is a basic way to explore structure in data with no labels needed.
# Preview relevant columns
print(df.columns)
print(df[['Age','Annual Income (k$)','Spending Score (1-100)']].describe())
K-Means vs. DBSCAN#
You may have heard of K-Means.
K-Means forms round, equally-sized clusters.
But what if groups are messy? DBSCAN to the rescue!
- DBSCAN stands for Density-Based Spatial Clustering of Applications with Noise.
- It finds clusters of any shape, even with outliers.
- You do not need to pick the number of clusters first.
# Visualize customers by incomes and spending
import matplotlib.pyplot as plt
plt.scatter(df['Annual Income (k$)'], df['Spending Score (1-100)'], alpha=0.6)
plt.xlabel('Annual Income (k$)')
plt.ylabel('Spending Score (1-100)')
plt.title('Mall Customers: Income vs. Spending')
plt.show()
What is density?#
In DBSCAN, 'density' means areas where lots of points are close together.
- A cluster has many points in each others neighborhoods.
- Points with few nearby neighbors may be called 'noise' (outliers).
DBSCAN needs two settings:
- eps: how close points must be to consider as neighbors.
- min_samples: how many points needed to make a cluster.
# Prepare data for clustering
X = df[['Annual Income (k$)', 'Spending Score (1-100)']].copy()
# Standardize features
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)
Fitting DBSCAN#
We will use sklearns DBSCAN to find clusters.
Let us start with some basic values for eps and min_samples.
Experimentation is normal! DBSCAN gives different results if you tweak eps.
# Run DBSCAN clustering
from sklearn.cluster import DBSCAN
dbscan = DBSCAN(eps=0.5, min_samples=5)
labels = dbscan.fit_predict(X_scaled)
print(pd.Series(labels).value_counts())
# Plot DBSCAN results
plt.figure(figsize=(7,5))
colors = ['red','green','blue','purple','orange','pink','gray']
for k in set(labels):
color = 'black' if k == -1 else colors[k%len(colors)]
class_member_mask = (labels == k)
plt.scatter(X[class_member_mask]['Annual Income (k$)'],
X[class_member_mask]['Spending Score (1-100)'],
c=color, label='Cluster '+str(k) if k!=-1 else 'Noise', alpha=0.5)
plt.xlabel('Annual Income (k$)')
plt.ylabel('Spending Score (1-100)')
plt.title('DBSCAN Clusters')
plt.legend()
plt.show()
# Change parameters to explore more
eps = float(input("Try a new eps (example 0.2 to 1.0): "))
min_samples = int(input("Try a new min_samples (example 3 to 10): "))
dbscan2 = DBSCAN(eps=eps, min_samples=min_samples)
labels2 = dbscan2.fit_predict(X_scaled)
plt.figure(figsize=(7,5))
for k in set(labels2):
color = 'black' if k == -1 else colors[k%len(colors)]
class_member_mask = (labels2 == k)
plt.scatter(X[class_member_mask]['Annual Income (k$)'],
X[class_member_mask]['Spending Score (1-100)'],
c=color, label='Cluster '+str(k) if k!=-1 else 'Noise', alpha=0.5)
plt.xlabel('Annual Income (k$)')
plt.ylabel('Spending Score (1-100)')
plt.title('DBSCAN: Your Settings')
plt.legend()
plt.show()
Understanding DBSCAN Labels#
Each point gets a label like 0, 1, 2 ... or -1.
-1 means the point is 'noise' it is not close to any cluster.
Inspecting and counting cluster sizes helps you tell if your clustering worked.
# Get a summary of each group
import numpy as np
for k in set(labels):
class_mask = (labels == k)
print(f"Cluster {k if k!=-1 else 'Noise'}: Count: {class_mask.sum()}")
print(' Mean income:', np.round(X[class_mask]['Annual Income (k$)'].mean(),2),
'Mean score:', np.round(X[class_mask]['Spending Score (1-100)'].mean(),2))
print('-----')
# Comparing DBSCAN to KMeans
from sklearn.cluster import KMeans
kmeans = KMeans(n_clusters=3, random_state=42)
k_labels = kmeans.fit_predict(X_scaled)
plt.figure(figsize=(7,5))
for k in set(k_labels):
class_mask = (k_labels == k)
plt.scatter(X[class_mask]['Annual Income (k$)'], X[class_mask]['Spending Score (1-100)'],
c=colors[k%len(colors)], label='KMeans '+str(k), alpha=0.5)
plt.xlabel('Annual Income (k$)')
plt.ylabel('Spending Score (1-100)')
plt.title('K-Means Clustering (for comparison)')
plt.legend()
plt.show()
Best practices for DBSCAN#
- Always scale your features first.
- Try several eps and min_samples values to get good results.
- Look at both the cluster plot and group summaries.
- Watch out for all-noise or too-large clusters adjust settings if you see these.
Remember: DBSCAN works best when clusters have dense, separate groups.
# Troubleshooting DBSCAN clustering
if (labels == -1).all():
print("All data points are marked as noise. Try increasing eps.")
elif len(set(labels)) == 1 and -1 not in set(labels):
print("Only one big cluster found. Try lowering eps or increasing min_samples.")
else:
print("DBSCAN found:", len(set(labels)) - (1 if -1 in set(labels) else 0), "clusters and", sum(labels==-1), "noise points.")
# Extra: Automatic parameter search using NearestNeighbors
from sklearn.neighbors import NearestNeighbors
neighbors = NearestNeighbors(n_neighbors=5)
neighbors_fit = neighbors.fit(X_scaled)
distances, indices = neighbors_fit.kneighbors(X_scaled)
distances = np.sort(distances[:,4])
plt.plot(distances)
plt.title('K-distance Graph (for eps guess)')
plt.xlabel('Points sorted by distance')
plt.ylabel('5th Nearest Neighbor distance')
plt.show()
Mini-Challenge: Cluster customer types#
Can you find parameter values for DBSCAN that give at least two clusters and less than 20% noise?
- Try different eps values.
- Try different min_samples.
- Look at the summary printout and plots.
This is good practice for when you try DBSCAN on your own data!
Recap#
Today you learned:
- What DBSCAN is and why it is different than KMeans.
- How to try, critique, and tune its clustering.
- That scaling data is important!
Keep experimenting data mining is a skill you build with practice.
Thanks for learning!#
If you enjoyed this lesson, please like, subscribe, and share.
Try using DBSCAN on your own dataset and let us know your results in the comments below!
Happy clustering see you next time!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



