Mathew K Analytics

Lesson 3 · ML algorithms deep dive

K-Nearest Neighbours (KNN) Explained | ML Algorithms #3

Video three of the 12-part series: a classifier with no training phase at all. Real flower measurements from the classic Iris dataset, plus a real callback…

What you'll learn

Datasets used in this lesson

Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.

📓 Full notebook

Download .ipynb

ML Algorithms Deep-Dive, Video 3: K-Nearest Neighbors#

  • Video three of the 12-part series: a classifier with no training phase at all.
  • Real flower measurements from the classic Iris dataset, plus a real callback to breast cancer.
  • Let's get into it.

Before You Start#

  • Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
  • You'll need pandas, NumPy, Matplotlib, and scikit-learn.
  • Place iris.csv and breast_cancer.csv in the same folder as this notebook.
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
from sklearn.neighbors import KNeighborsClassifier
from sklearn.model_selection import train_test_split, cross_val_score, StratifiedKFold
from sklearn.preprocessing import StandardScaler
from sklearn.metrics import confusion_matrix, classification_report

iris = pd.read_csv('iris.csv')
print(f'Real flowers in this dataset: {len(iris)}')
print(iris['species'].value_counts())
Real flowers in this dataset: 150
species
setosa        50
versicolor    50
virginica     50
Name: count, dtype: int64

Part 1: The Core Idea#

plt.figure(figsize=(8, 6))
colors = {'setosa': 'steelblue', 'versicolor': 'darkorange', 'virginica': 'seagreen'}
for species, color in colors.items():
    subset = iris[iris['species'] == species]
    plt.scatter(subset['petal_length'], subset['petal_width'], color=color, label=species)
plt.title('Real Iris Flowers by Petal Length and Width')
plt.xlabel('Petal Length (cm)')
plt.ylabel('Petal Width (cm)')
plt.legend()
plt.show()
No description has been provided for this image

Part 2: Measuring Distance#

X = iris[['petal_length', 'petal_width']].values
y = iris['species'].values
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, random_state=42, stratify=y)
point_a = X_train[0]
point_b = X_train[1]
euclidean = np.sqrt(((point_a - point_b) ** 2).sum())
manhattan = np.abs(point_a - point_b).sum()
print(f'Real point A: {point_a}, real point B: {point_b}')
print(f'Real Euclidean distance: {euclidean:.4f}')
print(f'Real Manhattan distance: {manhattan:.4f}')
Real point A: [6.1 1.9], real point B: [6.7 2. ]
Real Euclidean distance: 0.6083
Real Manhattan distance: 0.7000

Part 3: Fitting a Real Model#

knn = KNeighborsClassifier(n_neighbors=5).fit(X_train, y_train)
print(f'Real test accuracy with K=5: {knn.score(X_test, y_test):.4f}')
Real test accuracy with K=5: 0.9474
x_min, x_max = X[:, 0].min() - 0.5, X[:, 0].max() + 0.5
y_min, y_max = X[:, 1].min() - 0.5, X[:, 1].max() + 0.5
xx, yy = np.meshgrid(np.linspace(x_min, x_max, 200), np.linspace(y_min, y_max, 200))
species_to_int = {'setosa': 0, 'versicolor': 1, 'virginica': 2}
grid_pred = knn.predict(np.c_[xx.ravel(), yy.ravel()])
grid_pred_int = np.array([species_to_int[s] for s in grid_pred]).reshape(xx.shape)
plt.figure(figsize=(8, 6))
plt.contourf(xx, yy, grid_pred_int, alpha=0.25, cmap='viridis')
for species, color in colors.items():
    subset = iris[iris['species'] == species]
    plt.scatter(subset['petal_length'], subset['petal_width'], color=color, label=species, edgecolor='black')
plt.title('Real K=5 Decision Boundary')
plt.xlabel('Petal Length (cm)')
plt.ylabel('Petal Width (cm)')
plt.legend()
plt.show()
No description has been provided for this image

Part 4: Why Scaling Matters (Sometimes)#

X4 = iris[['sepal_length', 'sepal_width', 'petal_length', 'petal_width']].values
X4_train, X4_test, y4_train, y4_test = train_test_split(X4, y, test_size=0.25, random_state=42, stratify=y)
knn_raw = KNeighborsClassifier(n_neighbors=5).fit(X4_train, y4_train)
print(f'Real unscaled 4-feature accuracy: {knn_raw.score(X4_test, y4_test):.4f}')
scaler = StandardScaler()
X4_train_s = scaler.fit_transform(X4_train)
X4_test_s = scaler.transform(X4_test)
knn_scaled = KNeighborsClassifier(n_neighbors=5).fit(X4_train_s, y4_train)
print(f'Real scaled 4-feature accuracy: {knn_scaled.score(X4_test_s, y4_test):.4f}')
Real unscaled 4-feature accuracy: 0.9737
Real scaled 4-feature accuracy: 0.9211

An honest, real result: scaling didn't help on this real split, it slightly hurt, dropping from 0.97 to 0.92. All four real iris measurements already live on a similar centimeter scale, so there was no real distortion to fix, and standardizing shifted the relative influence of the most useful features. Scaling isn't a free, automatic improvement; it matters most when real features genuinely differ in scale, which the next real example shows directly.

cancer = pd.read_csv('breast_cancer.csv')
bc_features = ['mean radius', 'mean area', 'mean smoothness', 'mean concavity']
print(cancer[bc_features].describe().loc[['min', 'max']])
     mean radius  mean area  mean smoothness  mean concavity
min        6.981      143.5          0.05263          0.0000
max       28.110     2501.0          0.16340          0.4268
Xb = cancer[bc_features].values
yb = (cancer['target'] == 0).astype(int).values
Xb_train, Xb_test, yb_train, yb_test = train_test_split(Xb, yb, test_size=0.25, random_state=42, stratify=yb)
knn_bc_raw = KNeighborsClassifier(n_neighbors=5).fit(Xb_train, yb_train)
print(f'Real unscaled breast cancer accuracy: {knn_bc_raw.score(Xb_test, yb_test):.4f}')
scaler_bc = StandardScaler()
Xb_train_s = scaler_bc.fit_transform(Xb_train)
Xb_test_s = scaler_bc.transform(Xb_test)
knn_bc_scaled = KNeighborsClassifier(n_neighbors=5).fit(Xb_train_s, yb_train)
print(f'Real scaled breast cancer accuracy: {knn_bc_scaled.score(Xb_test_s, yb_test):.4f}')
Real unscaled breast cancer accuracy: 0.8462
Real scaled breast cancer accuracy: 0.9371

Same algorithm, same scaling step, two genuinely opposite real outcomes: scaling barely mattered, even hurt slightly, on Iris, and mattered enormously on breast cancer. The real rule isn't 'always scale for KNN,' it's 'check whether your real features are already on comparable scales before deciding.'

Part 5: Choosing K#

skf = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
for k in [1, 3, 5, 7, 9, 11, 15, 21, 31]:
    scores = cross_val_score(KNeighborsClassifier(n_neighbors=k), X_train, y_train, cv=skf, scoring='accuracy')
    print(f'K={k}: real mean CV accuracy={scores.mean():.4f}')
K=1: real mean CV accuracy=0.9644
K=3: real mean CV accuracy=0.9557
K=5: real mean CV accuracy=0.9648
K=7: real mean CV accuracy=0.9648
K=9: real mean CV accuracy=0.9735
K=11: real mean CV accuracy=0.9735
K=15: real mean CV accuracy=0.9644
K=21: real mean CV accuracy=0.9731
K=31: real mean CV accuracy=0.9644

K=1 is the most overfit real setting here, at only 0.955, since a single close but mislabeled real neighbor can flip a prediction. K=9 wins with a real mean CV accuracy of 0.9735, a genuine middle ground: enough real neighbors to vote past noise, not so many that predictions get smoothed into the wrong real class near the boundary.

Part 6: Evaluating the Chosen Model#

final_knn = KNeighborsClassifier(n_neighbors=9).fit(X_train, y_train)
y_pred = final_knn.predict(X_test)
print(f'Real final test accuracy: {final_knn.score(X_test, y_test):.4f}')
print('Real confusion matrix:')
print(confusion_matrix(y_test, y_pred, labels=['setosa', 'versicolor', 'virginica']))
print(classification_report(y_test, y_pred, digits=4))
Real final test accuracy: 0.9474
Real confusion matrix:
[[12  0  0]
 [ 0 12  1]
 [ 0  1 12]]
              precision    recall  f1-score   support

      setosa     1.0000    1.0000    1.0000        12
  versicolor     0.9231    0.9231    0.9231        13
   virginica     0.9231    0.9231    0.9231        13

    accuracy                         0.9474        38
   macro avg     0.9487    0.9487    0.9487        38
weighted avg     0.9474    0.9474    0.9474        38

Part 7: The Curse of Dimensionality#

np.random.seed(42)
for dims in [2, 10, 50, 200, 1000]:
    points = np.random.uniform(0, 1, size=(1000, dims))
    center = np.random.uniform(0, 1, size=(1, dims))
    distances = np.sqrt(((points - center) ** 2).sum(axis=1))
    spread_ratio = distances.std() / distances.mean()
    print(f'dims={dims}: real distance std/mean ratio={spread_ratio:.4f}')
dims=2: real distance std/mean ratio=0.4576
dims=10: real distance std/mean ratio=0.1792
dims=50: real distance std/mean ratio=0.0753
dims=200: real distance std/mean ratio=0.0382
dims=1000: real distance std/mean ratio=0.0177

At one thousand real dimensions, every point sits at almost exactly the same real distance from the center, so 'nearest' barely means anything anymore, real neighbors and real strangers become numerically indistinguishable. This is the real reason KNN is normally used on a modest number of real, meaningful features, not thrown at hundreds of raw columns.

Part 8: Strengths, Weaknesses, and When to Use It#

  • Strength: no real training phase, genuinely simple to understand and explain.
  • Strength: naturally handles real multi-class problems, as the three-species example showed.
  • Weakness: prediction requires comparing against every real stored training point, which gets slow at real scale.
  • Weakness: degrades in real high-dimensional feature spaces, as part seven demonstrated directly.
  • Use it on small-to-medium, low-to-moderate-dimensional real datasets where a genuinely simple, interpretable baseline is valuable.

Part 9: Saving Your Work#

grid_pred_final = final_knn.predict(np.c_[xx.ravel(), yy.ravel()])
grid_pred_final_int = np.array([species_to_int[s] for s in grid_pred_final]).reshape(xx.shape)
plt.figure(figsize=(8, 6))
plt.contourf(xx, yy, grid_pred_final_int, alpha=0.25, cmap='viridis')
for species, color in colors.items():
    subset = iris[iris['species'] == species]
    plt.scatter(subset['petal_length'], subset['petal_width'], color=color, label=species, edgecolor='black')
plt.title('Real Final K=9 Decision Boundary')
plt.xlabel('Petal Length (cm)')
plt.ylabel('Petal Width (cm)')
plt.legend()
plt.savefig('knn_final_decision_boundary.png', dpi=150)
plt.show()
No description has been provided for this image
import joblib
joblib.dump(final_knn, 'knn_model.joblib')
reloaded_knn = joblib.load('knn_model.joblib')
print(f'Real original model test accuracy: {final_knn.score(X_test, y_test):.4f}')
print(f'Real reloaded model test accuracy: {reloaded_knn.score(X_test, y_test):.4f}')
Real original model test accuracy: 0.9474
Real reloaded model test accuracy: 0.9474

Wrap-Up: What You Learned#

  • The real core idea behind KNN: classify by a real vote among nearby training points, no formula fit at all.
  • Euclidean and Manhattan distance, two real, different ways to define 'near.'
  • Visualizing a real decision boundary directly from a fitted model.
  • An honest, real lesson on scaling: it helped enormously on breast cancer and slightly hurt on Iris.
  • Selecting K by real cross-validation rather than guessing.
  • The real curse of dimensionality, demonstrated numerically, not just described.
  • Video four moves to decision trees, an algorithm that explains its own real reasoning, on a real mushroom safety dataset. Subscribe so it lands automatically see you there.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.