Mathew K Analytics

Lesson 9 · ML algorithms deep dive

K-Means Clustering in Python | ML Algorithms #9

Video nine of the 12-part series: the first algorithm in this series with no real labels at all. Real mall customer data: real age, real income, real…

What you'll learn

Datasets used in this lesson

Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.

📓 Full notebook

Download .ipynb

ML Algorithms Deep-Dive, Video 9: K-Means Clustering#

  • Video nine of the 12-part series: the first algorithm in this series with no real labels at all.
  • Real mall customer data: real age, real income, real spending behavior.
  • Let's get into it.

Before You Start#

  • Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
  • You'll need pandas, NumPy, Matplotlib, and scikit-learn.
  • Place mall_customers.csv in the same folder as this notebook.
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
from sklearn.cluster import KMeans
from sklearn.preprocessing import StandardScaler
from sklearn.metrics import silhouette_score

customers = pd.read_csv('mall_customers.csv')
print(f'Real customers in this dataset: {len(customers)}')
print(customers[['Age', 'Annual Income (k$)', 'Spending Score (1-100)']].describe())
Real customers in this dataset: 200
              Age  Annual Income (k$)  Spending Score (1-100)
count  200.000000          200.000000              200.000000
mean    38.850000           60.560000               50.200000
std     13.969007           26.264721               25.823522
min     18.000000           15.000000                1.000000
25%     28.750000           41.500000               34.750000
50%     36.000000           61.500000               50.000000
75%     49.000000           78.000000               73.000000
max     70.000000          137.000000               99.000000

Part 1: Unsupervised Learning: No Real Labels This Time#

plt.figure(figsize=(8, 6))
plt.scatter(customers['Annual Income (k$)'], customers['Spending Score (1-100)'], color='steelblue')
plt.title('Real Mall Customers: Income vs. Spending Score')
plt.xlabel('Annual Income (k$)')
plt.ylabel('Spending Score (1-100)')
plt.show()
No description has been provided for this image

Part 2: K-Means, Step by Step#

X = customers[['Annual Income (k$)', 'Spending Score (1-100)']].values
rng = np.random.RandomState(42)
initial_idx = rng.choice(len(X), size=3, replace=False)
centroids = X[initial_idx].astype(float)
print(f'Real starting centroids, picked from 3 random real customers:')
print(centroids)
Real starting centroids, picked from 3 random real customers:
[[60. 52.]
 [20. 79.]
 [30.  4.]]
def assign_clusters(points, centers):
    distances = np.sqrt(((points[:, None, :] - centers[None, :, :]) ** 2).sum(axis=2))
    return distances.argmin(axis=1), distances

labels, distances = assign_clusters(X, centroids)
inertia_round_0 = sum(distances[i, labels[i]] ** 2 for i in range(len(X)))
print(f'Real total squared distance after the first assignment: {inertia_round_0:.2f}')
Real total squared distance after the first assignment: 184114.00
updated_centroids = np.array([X[labels == k].mean(axis=0) for k in range(3)])
print('Real updated centroids after averaging their assigned real customers:')
print(updated_centroids)
labels_2, distances_2 = assign_clusters(X, updated_centroids)
inertia_round_1 = sum(distances_2[i, labels_2[i]] ** 2 for i in range(len(X)))
print(f'Real total squared distance after one full real iteration: {inertia_round_1:.2f}')
Real updated centroids after averaging their assigned real customers:
[[69.91946309 52.61744966]
 [25.72727273 79.36363636]
 [38.89655172 15.65517241]]
Real total squared distance after one full real iteration: 158118.90

Part 3: Choosing K: The Elbow Method#

inertias = []
for k in range(1, 11):
    model = KMeans(n_clusters=k, random_state=42, n_init=10).fit(X)
    inertias.append(model.inertia_)
    print(f'k={k}: real inertia={model.inertia_:.2f}')
k=1: real inertia=269981.28
k=2: real inertia=181363.60
k=3: real inertia=106348.37
k=4: real inertia=73679.79
k=5: real inertia=44448.46
k=6: real inertia=37233.81
k=7: real inertia=30241.34
k=8: real inertia=25036.42
k=9: real inertia=21916.79
k=10: real inertia=20072.07
plt.figure(figsize=(8, 5))
plt.plot(range(1, 11), inertias, marker='o', color='steelblue')
plt.title('Real Elbow Plot')
plt.xlabel('Number of Clusters (K)')
plt.ylabel('Inertia')
plt.show()
No description has been provided for this image

Part 4: Choosing K: The Silhouette Score#

for k in range(2, 9):
    model = KMeans(n_clusters=k, random_state=42, n_init=10).fit(X)
    score = silhouette_score(X, model.labels_)
    print(f'k={k}: real silhouette score={score:.4f}')
k=2: real silhouette score=0.2969
k=3: real silhouette score=0.4676
k=4: real silhouette score=0.4932
k=5: real silhouette score=0.5539
k=6: real silhouette score=0.5398
k=7: real silhouette score=0.5288
k=8: real silhouette score=0.4548

Part 5: Fitting the Real Final Model#

final_model = KMeans(n_clusters=5, random_state=42, n_init=10).fit(X)
customers['cluster'] = final_model.labels_
print(customers['cluster'].value_counts().sort_index())
cluster
0    81
1    39
2    22
3    35
4    23
Name: count, dtype: int64
plt.figure(figsize=(8, 6))
plt.scatter(customers['Annual Income (k$)'], customers['Spending Score (1-100)'], c=customers['cluster'], cmap='viridis')
plt.scatter(final_model.cluster_centers_[:, 0], final_model.cluster_centers_[:, 1], c='red', marker='X', s=200, label='Centroids')
plt.title('Real K=5 Mall Customer Segments')
plt.xlabel('Annual Income (k$)')
plt.ylabel('Spending Score (1-100)')
plt.legend()
plt.savefig('kmeans_final_clusters.png', dpi=150)
plt.show()
No description has been provided for this image

Part 6: Interpreting the Real Clusters#

cluster_summary = customers.groupby('cluster')[['Annual Income (k$)', 'Spending Score (1-100)']].mean()
cluster_summary['count'] = customers['cluster'].value_counts().sort_index()
print(cluster_summary)
         Annual Income (k$)  Spending Score (1-100)  count
cluster                                                   
0                 55.296296               49.518519     81
1                 86.538462               82.128205     39
2                 25.727273               79.363636     22
3                 88.200000               17.114286     35
4                 26.304348               20.913043     23

Part 7: Initialization Sensitivity#

for seed in [0, 1, 2, 3, 4]:
    single_run = KMeans(n_clusters=5, random_state=seed, n_init=1, init='random').fit(X)
    print(f'seed={seed}, single random init: real inertia={single_run.inertia_:.2f}')
seed=0, single random init: real inertia=44448.46
seed=1, single random init: real inertia=44448.46
seed=2, single random init: real inertia=44448.46
seed=3, single random init: real inertia=44448.46
seed=4, single random init: real inertia=89897.46

That real, worse-scoring seed is exactly the kind of outcome scikit-learn's default n_init of 10 real attempts, keeping only the real best, is designed to avoid. It's a small real detail easy to overlook, and a real reason to never trust a single K-Means run without it.

Part 8: Does Scaling Help Here?#

X3 = customers[['Age', 'Annual Income (k$)', 'Spending Score (1-100)']].values
raw_model = KMeans(n_clusters=5, random_state=42, n_init=10).fit(X3)
raw_silhouette = silhouette_score(X3, raw_model.labels_)
print(f'Real unscaled 3-feature silhouette: {raw_silhouette:.4f}')
scaler = StandardScaler()
X3_scaled = scaler.fit_transform(X3)
scaled_model = KMeans(n_clusters=5, random_state=42, n_init=10).fit(X3_scaled)
scaled_silhouette = silhouette_score(X3_scaled, scaled_model.labels_)
print(f'Real scaled 3-feature silhouette: {scaled_silhouette:.4f}')
Real unscaled 3-feature silhouette: 0.4443
Real scaled 3-feature silhouette: 0.4166

Part 9: Strengths, Weaknesses, and When to Use It#

  • Strength: fast and genuinely simple to run, even on real, larger datasets.
  • Strength: the elbow method and silhouette score give two independent, real ways to choose K.
  • Weakness: assumes real clusters are roughly round and similarly sized, genuinely different real shapes can fool it.
  • Weakness: sensitive to real initialization, as part seven demonstrated directly, always use multiple real restarts.
  • Use it for real customer segmentation or any exploratory task where genuine natural groupings, not predictions, are the goal.

Part 10: Saving Your Work#

import joblib
joblib.dump(final_model, 'kmeans_model.joblib')
reloaded_kmeans = joblib.load('kmeans_model.joblib')
print(f'Real original model inertia: {final_model.inertia_:.4f}')
print(f'Real reloaded model inertia: {reloaded_kmeans.inertia_:.4f}')
Real original model inertia: 44448.4554
Real reloaded model inertia: 44448.4554
customers[['CustomerID', 'Annual Income (k$)', 'Spending Score (1-100)', 'cluster']].to_csv('mall_customer_clusters.csv', index=False)
print('Real cluster assignments saved to mall_customer_clusters.csv')
Real cluster assignments saved to mall_customer_clusters.csv

Wrap-Up: What You Learned#

  • Unsupervised learning: finding genuine real structure with no real answer key at all.
  • K-Means' real assign-then-update loop, built by hand for one full real iteration.
  • Two independent real methods, the elbow plot and the silhouette score, agreeing on the same real K.
  • Translating real cluster centroids into a genuine, plain-language business interpretation.
  • A real, concrete demonstration of initialization sensitivity, and why multiple restarts matter.
  • Another honest, real case where scaling didn't actually help.
  • Video ten moves to hierarchical clustering, revisiting these same real mall customers with a genuinely different clustering approach. Subscribe so it lands automatically see you there.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.