Lesson 27 · Data Mining
Understanding 5 Key Cluster Evaluation Metrics for Effective Machine Learning Analysis
Clustering helps us discover groups within data. But how do we know if our groups are meaningful? This lesson will show you basic ways to measure and…
- CourseData Mining
- Lesson27 of 31
- Video15 min
- FormatJupyter notebook · 14 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbWeek 7: Cluster Evaluation Metrics and Interpretation#
Clustering helps us discover groups within data. But how do we know if our groups are meaningful?
This lesson will show you basic ways to measure and understand clustering results, step by step.
import warnings
import numpy as np
np.random.seed(42)
warnings.filterwarnings('ignore')
# We suppress warnings so outputs are cleaner.
# Data setup (Mall Customers Dataset)
import pandas as pd
url = 'https://gist.githubusercontent.com/pravalliyaram/5c05f43d2351249927b8a3f3cc3e5ecf/raw/Mall_Customers.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
About the Data#
Each row represents a customer and basic information like age, gender, income, and spending score.
We want to find groups of similar customers using clustering.
# Select only relevant numerical features for clustering
X = df[['Age', 'Annual Income (k$)', 'Spending Score (1-100)']]
print(X.head())
# Standardize our features (very important for clustering!)
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)
What is Clustering?#
Clustering tries to group data points so those in the same group are similar.
People use clustering in marketing, biology, and more.
# Fit KMeans clustering with k=3
from sklearn.cluster import KMeans
kmeans = KMeans(n_clusters=3, random_state=42)
labels = kmeans.fit_predict(X_scaled)
# Scatter plot to view clusters
import matplotlib.pyplot as plt
plt.figure(figsize=(6,4))
plt.scatter(X_scaled[:, 1], X_scaled[:, 2], c=labels, cmap='viridis', alpha=0.7)
plt.xlabel('Annual Income (scaled)')
plt.ylabel('Spending Score (scaled)')
plt.title('Cluster assignments')
plt.show()
How do we Know if Clusters are Good?#
Visuals help, but we need metrics to really judge clusters.
There are many ways: inertia, silhouette score, Davies-Bouldin, Calinski-Harabasz, and more.
# Inertia: sum of squared distances to nearest center (lower is better)
print('Inertia:', kmeans.inertia_)
# Silhouette score measures how similar a point is to its cluster vs others (-1=bad, 1=good)
from sklearn.metrics import silhouette_score
score = silhouette_score(X_scaled, labels)
print('Silhouette Score:', round(score,2))
# Davies-Bouldin Index: lower is better (measures overlap)
from sklearn.metrics import davies_bouldin_score
db_score = davies_bouldin_score(X_scaled, labels)
print('Davies-Bouldin Index:', round(db_score,2))
# Calinski-Harabasz Score: Higher is better
from sklearn.metrics import calinski_harabasz_score
ch_score = calinski_harabasz_score(X_scaled, labels)
print('Calinski-Harabasz Score:', int(ch_score))
# Try different k, compare silhouette scores
scores = []
for k in range(2,7):
km = KMeans(n_clusters=k, random_state=42)
labs = km.fit_predict(X_scaled)
score = silhouette_score(X_scaled, labs)
scores.append(round(score,2))
print('Silhouette scores for k=2 to k=6:', scores)
Interpreting Cluster Results#
Look for well-separated, compact clusters with good scores.
Remember: Metrics are just guidesreal meaning comes from knowing the business or system.
No metric can tell you if clusters truly matter for your goal.
# See what our clusters look like by average values
df['Cluster'] = labels
print(df.groupby('Cluster')[['Age', 'Annual Income (k$)', 'Spending Score (1-100)']].mean().round(1))
# User challenge: Try clustering with your own choice of k using input()
k = int(input('Choose a value for k (2-6): '))
custom_km = KMeans(n_clusters=k, random_state=42)
custom_labels = custom_km.fit_predict(X_scaled)
print('Silhouette Score:', round(silhouette_score(X_scaled, custom_labels), 2))
Common Pitfalls#
- Letting outliers affect clusters too much
- Using non-numeric data by accident
- Assuming metrics always mean business value
Fix these with good prep, checks, and business sense.
# Pro tip: Use elbow method to find best k visually
inertias = []
K = range(1, 8)
for k in K:
km = KMeans(n_clusters=k, random_state=42)
km.fit(X_scaled)
inertias.append(km.inertia_)
plt.plot(K, inertias, marker='o')
plt.xlabel('Number of clusters k')
plt.ylabel('Inertia')
plt.title('Elbow Method for optimal k')
plt.show()
Challenge:#
- Try different sets of features, like just age and income.
- Run clustering, then interpret cluster results.
- Plot their silhouette scores, compare, and explain in a sentence.
Share your results and favorite insights!
Recap#
- Cluster evaluation is about checking if groups are meaningful, not just splitting up data.
- Use multiple metrics like silhouette score and inertia.
- Always connect your results back to the real world goal.
With practice, this gets easier!
Thanks for Learning Cluster Evaluation!#
Try these techniques with your own datasets next.
If this was helpful, like, subscribe, and see the channel for more hands-on lessons.
See you in the next week!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



