Lesson 12 · Scikit-learn deep dive
Scikit-learn Tutorial #12: Clustering Metrics
Video twelve of the eighteen-part series: judging clusters when there's usually no ground truth. silhouettescore, the elbow method, and adjustedrandscore.…
- CourseScikit-learn deep dive
- Lesson12 of 18
- Video14 min
- FormatJupyter notebook · 10 code cells
What you'll learn
- The Problem - Clustering Usually Has No Ground Truth
- silhouettescore Basics
- silhouettesamples - Per-Point Scores
- The Elbow Method - inertia Across k
- Choosing k with Silhouette Across k
- adjustedrandscore - When True Labels ARE Known
- normalizedmutualinfoscore - an Alternative External Metric
- calinskiharabaszscore and daviesbouldinscore
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbScikit-learn Deep-Dive, Video 12: Clustering Metrics#
- Video twelve of the eighteen-part series: judging clusters when there's usually no ground truth.
- silhouette_score, the elbow method, and adjusted_rand_score.
- Let's get into it.
Part 1: The Problem - Clustering Usually Has No Ground Truth#
from sklearn.datasets import load_iris
from sklearn.cluster import KMeans
X, y = load_iris(return_X_y=True)
km = KMeans(n_clusters=3, random_state=42, n_init=10)
cluster_labels = km.fit_predict(X)
print(cluster_labels[:10])
Part 2: silhouette_score Basics#
from sklearn.metrics import silhouette_score
sil = silhouette_score(X, cluster_labels)
print(round(sil, 3))
Part 3: silhouette_samples - Per-Point Scores#
from sklearn.metrics import silhouette_samples
sample_scores = silhouette_samples(X, cluster_labels)
print(sample_scores[:10].round(3))
print((sample_scores < 0).sum())
Part 4: The Elbow Method - inertia_ Across k#
inertias = []
k_range = range(1, 8)
for k in k_range:
km_k = KMeans(n_clusters=k, random_state=42, n_init=10).fit(X)
inertias.append(round(km_k.inertia_, 1))
print(list(zip(k_range, inertias)))
Part 5: Choosing k with Silhouette Across k#
sil_scores = {}
for k in range(2, 8):
labels_k = KMeans(n_clusters=k, random_state=42, n_init=10).fit_predict(X)
sil_scores[k] = round(silhouette_score(X, labels_k), 3)
print(sil_scores)
print(max(sil_scores, key=sil_scores.get))
Part 6: adjusted_rand_score - When True Labels ARE Known#
import numpy as np
from sklearn.metrics import adjusted_rand_score
ari = adjusted_rand_score(y, cluster_labels)
print(round(ari, 3))
random_labels = np.random.RandomState(0).randint(0, 3, size=len(y))
print(round(adjusted_rand_score(y, random_labels), 3))
Part 7: normalized_mutual_info_score - an Alternative External Metric#
from sklearn.metrics import normalized_mutual_info_score
nmi = normalized_mutual_info_score(y, cluster_labels)
print(round(nmi, 3))
Part 8: calinski_harabasz_score and davies_bouldin_score#
from sklearn.metrics import calinski_harabasz_score, davies_bouldin_score
ch = calinski_harabasz_score(X, cluster_labels)
db = davies_bouldin_score(X, cluster_labels)
print(round(ch, 1))
print(round(db, 3))
Part 9: Comparing KMeans vs AgglomerativeClustering with Metrics#
from sklearn.cluster import AgglomerativeClustering
agg = AgglomerativeClustering(n_clusters=3)
agg_labels = agg.fit_predict(X)
print('KMeans silhouette:', round(silhouette_score(X, cluster_labels), 3))
print('Agglomerative silhouette:', round(silhouette_score(X, agg_labels), 3))
print('KMeans ARI:', round(adjusted_rand_score(y, cluster_labels), 3))
print('Agglomerative ARI:', round(adjusted_rand_score(y, agg_labels), 3))
Part 10: A Real Pattern - a Reusable find_best_k Function#
def find_best_k(X, k_range=range(2, 8)):
best_k, best_score = None, -1
for k in k_range:
labels = KMeans(n_clusters=k, random_state=42, n_init=10).fit_predict(X)
score = silhouette_score(X, labels)
if score > best_score:
best_k, best_score = k, score
return best_k, round(best_score, 3)
print(find_best_k(X))
Wrap-Up: What You Learned#
- Clustering usually has no ground truth; internal metrics judge quality from the data and cluster labels alone.
- silhouette_score measures cohesion versus separation, ranging from negative one to one, higher is better.
- silhouette_samples gives per-point scores, flagging possibly misassigned points.
- The elbow method looks at inertia_ across k, choosing the point where more clusters stop helping much.
- Silhouette score across a range of k gives a second, independent signal for choosing k.
- adjusted_rand_score compares cluster assignments to true labels when they exist, correcting for chance agreement.
- normalized_mutual_info_score is an alternative external metric measuring shared information between label sets.
- calinski_harabasz_score (higher is better) and davies_bouldin_score (lower is better) are further internal metrics.
- The same metric suite applies across different clustering algorithms for a fair, apples-to-apples comparison.
- That wraps up clustering metrics. Next up: Feature Selection - SelectKBest, RFE, and feature importances.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



