Lesson 9 · ML algorithms deep dive
K-Means Clustering in Python | ML Algorithms #9
Video nine of the 12-part series: the first algorithm in this series with no real labels at all. Real mall customer data: real age, real income, real…
- CourseML algorithms deep dive
- Lesson9 of 12
- Video19 min
- FormatJupyter notebook · 15 code cells
- Data1 dataset
What you'll learn
Datasets used in this lesson
Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.
- mall_customers.csv4.3 KB
📓 Full notebook
Download .ipynbML Algorithms Deep-Dive, Video 9: K-Means Clustering#
- Video nine of the 12-part series: the first algorithm in this series with no real labels at all.
- Real mall customer data: real age, real income, real spending behavior.
- Let's get into it.
Before You Start#
- Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
- You'll need pandas, NumPy, Matplotlib, and scikit-learn.
- Place mall_customers.csv in the same folder as this notebook.
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
from sklearn.cluster import KMeans
from sklearn.preprocessing import StandardScaler
from sklearn.metrics import silhouette_score
customers = pd.read_csv('mall_customers.csv')
print(f'Real customers in this dataset: {len(customers)}')
print(customers[['Age', 'Annual Income (k$)', 'Spending Score (1-100)']].describe())
Part 1: Unsupervised Learning: No Real Labels This Time#
plt.figure(figsize=(8, 6))
plt.scatter(customers['Annual Income (k$)'], customers['Spending Score (1-100)'], color='steelblue')
plt.title('Real Mall Customers: Income vs. Spending Score')
plt.xlabel('Annual Income (k$)')
plt.ylabel('Spending Score (1-100)')
plt.show()
Part 2: K-Means, Step by Step#
X = customers[['Annual Income (k$)', 'Spending Score (1-100)']].values
rng = np.random.RandomState(42)
initial_idx = rng.choice(len(X), size=3, replace=False)
centroids = X[initial_idx].astype(float)
print(f'Real starting centroids, picked from 3 random real customers:')
print(centroids)
def assign_clusters(points, centers):
distances = np.sqrt(((points[:, None, :] - centers[None, :, :]) ** 2).sum(axis=2))
return distances.argmin(axis=1), distances
labels, distances = assign_clusters(X, centroids)
inertia_round_0 = sum(distances[i, labels[i]] ** 2 for i in range(len(X)))
print(f'Real total squared distance after the first assignment: {inertia_round_0:.2f}')
updated_centroids = np.array([X[labels == k].mean(axis=0) for k in range(3)])
print('Real updated centroids after averaging their assigned real customers:')
print(updated_centroids)
labels_2, distances_2 = assign_clusters(X, updated_centroids)
inertia_round_1 = sum(distances_2[i, labels_2[i]] ** 2 for i in range(len(X)))
print(f'Real total squared distance after one full real iteration: {inertia_round_1:.2f}')
Part 3: Choosing K: The Elbow Method#
inertias = []
for k in range(1, 11):
model = KMeans(n_clusters=k, random_state=42, n_init=10).fit(X)
inertias.append(model.inertia_)
print(f'k={k}: real inertia={model.inertia_:.2f}')
plt.figure(figsize=(8, 5))
plt.plot(range(1, 11), inertias, marker='o', color='steelblue')
plt.title('Real Elbow Plot')
plt.xlabel('Number of Clusters (K)')
plt.ylabel('Inertia')
plt.show()
Part 4: Choosing K: The Silhouette Score#
for k in range(2, 9):
model = KMeans(n_clusters=k, random_state=42, n_init=10).fit(X)
score = silhouette_score(X, model.labels_)
print(f'k={k}: real silhouette score={score:.4f}')
Part 5: Fitting the Real Final Model#
final_model = KMeans(n_clusters=5, random_state=42, n_init=10).fit(X)
customers['cluster'] = final_model.labels_
print(customers['cluster'].value_counts().sort_index())
plt.figure(figsize=(8, 6))
plt.scatter(customers['Annual Income (k$)'], customers['Spending Score (1-100)'], c=customers['cluster'], cmap='viridis')
plt.scatter(final_model.cluster_centers_[:, 0], final_model.cluster_centers_[:, 1], c='red', marker='X', s=200, label='Centroids')
plt.title('Real K=5 Mall Customer Segments')
plt.xlabel('Annual Income (k$)')
plt.ylabel('Spending Score (1-100)')
plt.legend()
plt.savefig('kmeans_final_clusters.png', dpi=150)
plt.show()
Part 6: Interpreting the Real Clusters#
cluster_summary = customers.groupby('cluster')[['Annual Income (k$)', 'Spending Score (1-100)']].mean()
cluster_summary['count'] = customers['cluster'].value_counts().sort_index()
print(cluster_summary)
Part 7: Initialization Sensitivity#
for seed in [0, 1, 2, 3, 4]:
single_run = KMeans(n_clusters=5, random_state=seed, n_init=1, init='random').fit(X)
print(f'seed={seed}, single random init: real inertia={single_run.inertia_:.2f}')
That real, worse-scoring seed is exactly the kind of outcome scikit-learn's default n_init of 10 real attempts, keeping only the real best, is designed to avoid. It's a small real detail easy to overlook, and a real reason to never trust a single K-Means run without it.
Part 8: Does Scaling Help Here?#
X3 = customers[['Age', 'Annual Income (k$)', 'Spending Score (1-100)']].values
raw_model = KMeans(n_clusters=5, random_state=42, n_init=10).fit(X3)
raw_silhouette = silhouette_score(X3, raw_model.labels_)
print(f'Real unscaled 3-feature silhouette: {raw_silhouette:.4f}')
scaler = StandardScaler()
X3_scaled = scaler.fit_transform(X3)
scaled_model = KMeans(n_clusters=5, random_state=42, n_init=10).fit(X3_scaled)
scaled_silhouette = silhouette_score(X3_scaled, scaled_model.labels_)
print(f'Real scaled 3-feature silhouette: {scaled_silhouette:.4f}')
Part 9: Strengths, Weaknesses, and When to Use It#
- Strength: fast and genuinely simple to run, even on real, larger datasets.
- Strength: the elbow method and silhouette score give two independent, real ways to choose K.
- Weakness: assumes real clusters are roughly round and similarly sized, genuinely different real shapes can fool it.
- Weakness: sensitive to real initialization, as part seven demonstrated directly, always use multiple real restarts.
- Use it for real customer segmentation or any exploratory task where genuine natural groupings, not predictions, are the goal.
Part 10: Saving Your Work#
import joblib
joblib.dump(final_model, 'kmeans_model.joblib')
reloaded_kmeans = joblib.load('kmeans_model.joblib')
print(f'Real original model inertia: {final_model.inertia_:.4f}')
print(f'Real reloaded model inertia: {reloaded_kmeans.inertia_:.4f}')
customers[['CustomerID', 'Annual Income (k$)', 'Spending Score (1-100)', 'cluster']].to_csv('mall_customer_clusters.csv', index=False)
print('Real cluster assignments saved to mall_customer_clusters.csv')
Wrap-Up: What You Learned#
- Unsupervised learning: finding genuine real structure with no real answer key at all.
- K-Means' real assign-then-update loop, built by hand for one full real iteration.
- Two independent real methods, the elbow plot and the silhouette score, agreeing on the same real K.
- Translating real cluster centroids into a genuine, plain-language business interpretation.
- A real, concrete demonstration of initialization sensitivity, and why multiple restarts matter.
- Another honest, real case where scaling didn't actually help.
- Video ten moves to hierarchical clustering, revisiting these same real mall customers with a genuinely different clustering approach. Subscribe so it lands automatically see you there.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



