Mathew K Analytics

Lesson 12 · Scikit-learn deep dive

Scikit-learn Tutorial #12: Clustering Metrics

Video twelve of the eighteen-part series: judging clusters when there's usually no ground truth. silhouettescore, the elbow method, and adjustedrandscore.…

⬇ Download notebookOpen in Colab ↗

📓 Full notebook

Download .ipynb

Scikit-learn Deep-Dive, Video 12: Clustering Metrics#

  • Video twelve of the eighteen-part series: judging clusters when there's usually no ground truth.
  • silhouette_score, the elbow method, and adjusted_rand_score.
  • Let's get into it.

Part 1: The Problem - Clustering Usually Has No Ground Truth#

from sklearn.datasets import load_iris
from sklearn.cluster import KMeans
X, y = load_iris(return_X_y=True)
km = KMeans(n_clusters=3, random_state=42, n_init=10)
cluster_labels = km.fit_predict(X)
print(cluster_labels[:10])
[1 1 1 1 1 1 1 1 1 1]

Part 2: silhouette_score Basics#

from sklearn.metrics import silhouette_score
sil = silhouette_score(X, cluster_labels)
print(round(sil, 3))
0.553

Part 3: silhouette_samples - Per-Point Scores#

from sklearn.metrics import silhouette_samples
sample_scores = silhouette_samples(X, cluster_labels)
print(sample_scores[:10].round(3))
print((sample_scores < 0).sum())
[0.853 0.815 0.829 0.805 0.849 0.748 0.822 0.854 0.752 0.825]
0

Part 4: The Elbow Method - inertia_ Across k#

inertias = []
k_range = range(1, 8)
for k in k_range:
    km_k = KMeans(n_clusters=k, random_state=42, n_init=10).fit(X)
    inertias.append(round(km_k.inertia_, 1))
print(list(zip(k_range, inertias)))
[(1, 681.4), (2, 152.3), (3, 78.9), (4, 57.2), (5, 46.5), (6, 39.0), (7, 34.3)]

Part 5: Choosing k with Silhouette Across k#

sil_scores = {}
for k in range(2, 8):
    labels_k = KMeans(n_clusters=k, random_state=42, n_init=10).fit_predict(X)
    sil_scores[k] = round(silhouette_score(X, labels_k), 3)
print(sil_scores)
print(max(sil_scores, key=sil_scores.get))
{2: 0.681, 3: 0.553, 4: 0.498, 5: 0.491, 6: 0.365, 7: 0.354}
2

Part 6: adjusted_rand_score - When True Labels ARE Known#

import numpy as np
from sklearn.metrics import adjusted_rand_score
ari = adjusted_rand_score(y, cluster_labels)
print(round(ari, 3))
random_labels = np.random.RandomState(0).randint(0, 3, size=len(y))
print(round(adjusted_rand_score(y, random_labels), 3))
0.73
0.01

Part 7: normalized_mutual_info_score - an Alternative External Metric#

from sklearn.metrics import normalized_mutual_info_score
nmi = normalized_mutual_info_score(y, cluster_labels)
print(round(nmi, 3))
0.758

Part 8: calinski_harabasz_score and davies_bouldin_score#

from sklearn.metrics import calinski_harabasz_score, davies_bouldin_score
ch = calinski_harabasz_score(X, cluster_labels)
db = davies_bouldin_score(X, cluster_labels)
print(round(ch, 1))
print(round(db, 3))
561.6
0.662

Part 9: Comparing KMeans vs AgglomerativeClustering with Metrics#

from sklearn.cluster import AgglomerativeClustering
agg = AgglomerativeClustering(n_clusters=3)
agg_labels = agg.fit_predict(X)
print('KMeans silhouette:', round(silhouette_score(X, cluster_labels), 3))
print('Agglomerative silhouette:', round(silhouette_score(X, agg_labels), 3))
print('KMeans ARI:', round(adjusted_rand_score(y, cluster_labels), 3))
print('Agglomerative ARI:', round(adjusted_rand_score(y, agg_labels), 3))
KMeans silhouette: 0.553
Agglomerative silhouette: 0.554
KMeans ARI: 0.73
Agglomerative ARI: 0.731

Part 10: A Real Pattern - a Reusable find_best_k Function#

def find_best_k(X, k_range=range(2, 8)):
    best_k, best_score = None, -1
    for k in k_range:
        labels = KMeans(n_clusters=k, random_state=42, n_init=10).fit_predict(X)
        score = silhouette_score(X, labels)
        if score > best_score:
            best_k, best_score = k, score
    return best_k, round(best_score, 3)
print(find_best_k(X))
(2, 0.681)

Wrap-Up: What You Learned#

  • Clustering usually has no ground truth; internal metrics judge quality from the data and cluster labels alone.
  • silhouette_score measures cohesion versus separation, ranging from negative one to one, higher is better.
  • silhouette_samples gives per-point scores, flagging possibly misassigned points.
  • The elbow method looks at inertia_ across k, choosing the point where more clusters stop helping much.
  • Silhouette score across a range of k gives a second, independent signal for choosing k.
  • adjusted_rand_score compares cluster assignments to true labels when they exist, correcting for chance agreement.
  • normalized_mutual_info_score is an alternative external metric measuring shared information between label sets.
  • calinski_harabasz_score (higher is better) and davies_bouldin_score (lower is better) are further internal metrics.
  • The same metric suite applies across different clustering algorithms for a fair, apples-to-apples comparison.
  • That wraps up clustering metrics. Next up: Feature Selection - SelectKBest, RFE, and feature importances.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.