Lesson 3 · ML algorithms deep dive
K-Nearest Neighbours (KNN) Explained | ML Algorithms #3
Video three of the 12-part series: a classifier with no training phase at all. Real flower measurements from the classic Iris dataset, plus a real callback…
- CourseML algorithms deep dive
- Lesson3 of 12
- Video20 min
- FormatJupyter notebook · 13 code cells
- Data2 datasets
What you'll learn
Datasets used in this lesson
Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.
- iris.csv3.8 KB
- breast_cancer.csv118.5 KB
📓 Full notebook
Download .ipynbML Algorithms Deep-Dive, Video 3: K-Nearest Neighbors#
- Video three of the 12-part series: a classifier with no training phase at all.
- Real flower measurements from the classic Iris dataset, plus a real callback to breast cancer.
- Let's get into it.
Before You Start#
- Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
- You'll need pandas, NumPy, Matplotlib, and scikit-learn.
- Place iris.csv and breast_cancer.csv in the same folder as this notebook.
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
from sklearn.neighbors import KNeighborsClassifier
from sklearn.model_selection import train_test_split, cross_val_score, StratifiedKFold
from sklearn.preprocessing import StandardScaler
from sklearn.metrics import confusion_matrix, classification_report
iris = pd.read_csv('iris.csv')
print(f'Real flowers in this dataset: {len(iris)}')
print(iris['species'].value_counts())
Part 1: The Core Idea#
plt.figure(figsize=(8, 6))
colors = {'setosa': 'steelblue', 'versicolor': 'darkorange', 'virginica': 'seagreen'}
for species, color in colors.items():
subset = iris[iris['species'] == species]
plt.scatter(subset['petal_length'], subset['petal_width'], color=color, label=species)
plt.title('Real Iris Flowers by Petal Length and Width')
plt.xlabel('Petal Length (cm)')
plt.ylabel('Petal Width (cm)')
plt.legend()
plt.show()
Part 2: Measuring Distance#
X = iris[['petal_length', 'petal_width']].values
y = iris['species'].values
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, random_state=42, stratify=y)
point_a = X_train[0]
point_b = X_train[1]
euclidean = np.sqrt(((point_a - point_b) ** 2).sum())
manhattan = np.abs(point_a - point_b).sum()
print(f'Real point A: {point_a}, real point B: {point_b}')
print(f'Real Euclidean distance: {euclidean:.4f}')
print(f'Real Manhattan distance: {manhattan:.4f}')
Part 3: Fitting a Real Model#
knn = KNeighborsClassifier(n_neighbors=5).fit(X_train, y_train)
print(f'Real test accuracy with K=5: {knn.score(X_test, y_test):.4f}')
x_min, x_max = X[:, 0].min() - 0.5, X[:, 0].max() + 0.5
y_min, y_max = X[:, 1].min() - 0.5, X[:, 1].max() + 0.5
xx, yy = np.meshgrid(np.linspace(x_min, x_max, 200), np.linspace(y_min, y_max, 200))
species_to_int = {'setosa': 0, 'versicolor': 1, 'virginica': 2}
grid_pred = knn.predict(np.c_[xx.ravel(), yy.ravel()])
grid_pred_int = np.array([species_to_int[s] for s in grid_pred]).reshape(xx.shape)
plt.figure(figsize=(8, 6))
plt.contourf(xx, yy, grid_pred_int, alpha=0.25, cmap='viridis')
for species, color in colors.items():
subset = iris[iris['species'] == species]
plt.scatter(subset['petal_length'], subset['petal_width'], color=color, label=species, edgecolor='black')
plt.title('Real K=5 Decision Boundary')
plt.xlabel('Petal Length (cm)')
plt.ylabel('Petal Width (cm)')
plt.legend()
plt.show()
Part 4: Why Scaling Matters (Sometimes)#
X4 = iris[['sepal_length', 'sepal_width', 'petal_length', 'petal_width']].values
X4_train, X4_test, y4_train, y4_test = train_test_split(X4, y, test_size=0.25, random_state=42, stratify=y)
knn_raw = KNeighborsClassifier(n_neighbors=5).fit(X4_train, y4_train)
print(f'Real unscaled 4-feature accuracy: {knn_raw.score(X4_test, y4_test):.4f}')
scaler = StandardScaler()
X4_train_s = scaler.fit_transform(X4_train)
X4_test_s = scaler.transform(X4_test)
knn_scaled = KNeighborsClassifier(n_neighbors=5).fit(X4_train_s, y4_train)
print(f'Real scaled 4-feature accuracy: {knn_scaled.score(X4_test_s, y4_test):.4f}')
An honest, real result: scaling didn't help on this real split, it slightly hurt, dropping from 0.97 to 0.92. All four real iris measurements already live on a similar centimeter scale, so there was no real distortion to fix, and standardizing shifted the relative influence of the most useful features. Scaling isn't a free, automatic improvement; it matters most when real features genuinely differ in scale, which the next real example shows directly.
cancer = pd.read_csv('breast_cancer.csv')
bc_features = ['mean radius', 'mean area', 'mean smoothness', 'mean concavity']
print(cancer[bc_features].describe().loc[['min', 'max']])
Xb = cancer[bc_features].values
yb = (cancer['target'] == 0).astype(int).values
Xb_train, Xb_test, yb_train, yb_test = train_test_split(Xb, yb, test_size=0.25, random_state=42, stratify=yb)
knn_bc_raw = KNeighborsClassifier(n_neighbors=5).fit(Xb_train, yb_train)
print(f'Real unscaled breast cancer accuracy: {knn_bc_raw.score(Xb_test, yb_test):.4f}')
scaler_bc = StandardScaler()
Xb_train_s = scaler_bc.fit_transform(Xb_train)
Xb_test_s = scaler_bc.transform(Xb_test)
knn_bc_scaled = KNeighborsClassifier(n_neighbors=5).fit(Xb_train_s, yb_train)
print(f'Real scaled breast cancer accuracy: {knn_bc_scaled.score(Xb_test_s, yb_test):.4f}')
Same algorithm, same scaling step, two genuinely opposite real outcomes: scaling barely mattered, even hurt slightly, on Iris, and mattered enormously on breast cancer. The real rule isn't 'always scale for KNN,' it's 'check whether your real features are already on comparable scales before deciding.'
Part 5: Choosing K#
skf = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
for k in [1, 3, 5, 7, 9, 11, 15, 21, 31]:
scores = cross_val_score(KNeighborsClassifier(n_neighbors=k), X_train, y_train, cv=skf, scoring='accuracy')
print(f'K={k}: real mean CV accuracy={scores.mean():.4f}')
K=1 is the most overfit real setting here, at only 0.955, since a single close but mislabeled real neighbor can flip a prediction. K=9 wins with a real mean CV accuracy of 0.9735, a genuine middle ground: enough real neighbors to vote past noise, not so many that predictions get smoothed into the wrong real class near the boundary.
Part 6: Evaluating the Chosen Model#
final_knn = KNeighborsClassifier(n_neighbors=9).fit(X_train, y_train)
y_pred = final_knn.predict(X_test)
print(f'Real final test accuracy: {final_knn.score(X_test, y_test):.4f}')
print('Real confusion matrix:')
print(confusion_matrix(y_test, y_pred, labels=['setosa', 'versicolor', 'virginica']))
print(classification_report(y_test, y_pred, digits=4))
Part 7: The Curse of Dimensionality#
np.random.seed(42)
for dims in [2, 10, 50, 200, 1000]:
points = np.random.uniform(0, 1, size=(1000, dims))
center = np.random.uniform(0, 1, size=(1, dims))
distances = np.sqrt(((points - center) ** 2).sum(axis=1))
spread_ratio = distances.std() / distances.mean()
print(f'dims={dims}: real distance std/mean ratio={spread_ratio:.4f}')
At one thousand real dimensions, every point sits at almost exactly the same real distance from the center, so 'nearest' barely means anything anymore, real neighbors and real strangers become numerically indistinguishable. This is the real reason KNN is normally used on a modest number of real, meaningful features, not thrown at hundreds of raw columns.
Part 8: Strengths, Weaknesses, and When to Use It#
- Strength: no real training phase, genuinely simple to understand and explain.
- Strength: naturally handles real multi-class problems, as the three-species example showed.
- Weakness: prediction requires comparing against every real stored training point, which gets slow at real scale.
- Weakness: degrades in real high-dimensional feature spaces, as part seven demonstrated directly.
- Use it on small-to-medium, low-to-moderate-dimensional real datasets where a genuinely simple, interpretable baseline is valuable.
Part 9: Saving Your Work#
grid_pred_final = final_knn.predict(np.c_[xx.ravel(), yy.ravel()])
grid_pred_final_int = np.array([species_to_int[s] for s in grid_pred_final]).reshape(xx.shape)
plt.figure(figsize=(8, 6))
plt.contourf(xx, yy, grid_pred_final_int, alpha=0.25, cmap='viridis')
for species, color in colors.items():
subset = iris[iris['species'] == species]
plt.scatter(subset['petal_length'], subset['petal_width'], color=color, label=species, edgecolor='black')
plt.title('Real Final K=9 Decision Boundary')
plt.xlabel('Petal Length (cm)')
plt.ylabel('Petal Width (cm)')
plt.legend()
plt.savefig('knn_final_decision_boundary.png', dpi=150)
plt.show()
import joblib
joblib.dump(final_knn, 'knn_model.joblib')
reloaded_knn = joblib.load('knn_model.joblib')
print(f'Real original model test accuracy: {final_knn.score(X_test, y_test):.4f}')
print(f'Real reloaded model test accuracy: {reloaded_knn.score(X_test, y_test):.4f}')
Wrap-Up: What You Learned#
- The real core idea behind KNN: classify by a real vote among nearby training points, no formula fit at all.
- Euclidean and Manhattan distance, two real, different ways to define 'near.'
- Visualizing a real decision boundary directly from a fitted model.
- An honest, real lesson on scaling: it helped enormously on breast cancer and slightly hurt on Iris.
- Selecting K by real cross-validation rather than guessing.
- The real curse of dimensionality, demonstrated numerically, not just described.
- Video four moves to decision trees, an algorithm that explains its own real reasoning, on a real mushroom safety dataset. Subscribe so it lands automatically see you there.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



