Mathew K Analytics

Lesson 44 · Python Fundamentals

Understanding K Nearest Neighbors (KNN) in Python: A Step-by-Step Guide for Classification

In this lesson, we will learn how to use Python to find patterns and make predictions with a simple but powerful machine learning technique called KNN. You…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb
 

Welcome to K Nearest Neighbors (KNN) in Python#

In this lesson, we will learn how to use Python to find patterns and make predictions with a simple but powerful machine learning technique called KNN.

You will see how to work with real data, step by step.

No experience needed!

# Let's start with a quick setup.
import warnings
warnings.filterwarnings("ignore")

print("Ready to get started!")
Ready to get started!

What is KNN?#

KNN is a way for computers to recognize patterns by comparing things.

It asks: What are the most similar examples to this new one?

It is useful for classifying things (like sorting emails as spam or not spam!).

# Let's try some basic math first!
a = 9
b = 16
result = (a + b) / 2
print("Average:", result)
Average: 12.5
# Now let's use a list.
numbers = [10, 20, 30, 40, 50]
print("First number:", numbers[0])
print("Last number:", numbers[-1])
First number: 10
Last number: 50

How Does KNN Work?#

KNN finds the "K" closest neighbors to any point.

It counts which category shows up the most nearby.

Then, it guesses that your new data belongs to that category.

# Let's measure distance by hand with two numbers.
x1 = 7
x2 = 17
distance = abs(x1 - x2)
print("Distance:", distance)
Distance: 10
# Now let's use two features, like height and weight.
person_A = [170, 65]
person_B = [160, 60]
diff_height = person_A[0] - person_B[0]
diff_weight = person_A[1] - person_B[1]
euclidean_distance = (diff_height ** 2 + diff_weight ** 2) ** 0.5
print("Distance between A and B:", euclidean_distance)
Distance between A and B: 11.180339887498949

Let's Work With Real Data!#

We will use the famous Iris flower dataset.

We want to predict the species of a flower from some measurements.

# Data setup
from sklearn.datasets import load_iris
import pandas as pd

iris = load_iris(as_frame=True)
df = iris['frame']
print('Shape:', df.shape)
df.head()
Shape: (150, 5)
sepal length (cm) sepal width (cm) petal length (cm) petal width (cm) target
0 5.1 3.5 1.4 0.2 0
1 4.9 3.0 1.4 0.2 0
2 4.7 3.2 1.3 0.2 0
3 4.6 3.1 1.5 0.2 0
4 5.0 3.6 1.4 0.2 0
# Check the target distribution
df['target'].value_counts()
target
0    50
1    50
2    50
Name: count, dtype: int64
# Split the data into a training set and a test set
from sklearn.model_selection import train_test_split

X = df.drop('target', axis=1)
y = df['target']

X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
print('Training samples:', len(X_train))
print('Test samples:', len(X_test))
Training samples: 120
Test samples: 30
# Train a KNN model
from sklearn.neighbors import KNeighborsClassifier

knn = KNeighborsClassifier(n_neighbors=3)
knn.fit(X_train, y_train)

print('KNN model trained!')
KNN model trained!
# Predict on test data
y_pred = knn.predict(X_test)
print('Predicted labels:', y_pred[:10])
Predicted labels: [1 0 2 1 1 0 1 2 1 1]
# Measure accuracy
from sklearn.metrics import accuracy_score

accuracy = accuracy_score(y_test, y_pred)
print("Accuracy:", accuracy)
Accuracy: 1.0
# Try predicting a new flower using input()
values = []
feature_names = list(X.columns)
for name in feature_names:
    val = float(input(f"Enter {name}: "))
    values.append(val)
prediction = knn.predict([values])[0]
print("Predicted species:", iris['target_names'][prediction])
Predicted species: setosa
 
# What if we pick a bad value for k?
knn_bad = KNeighborsClassifier(n_neighbors=100)
knn_bad.fit(X_train, y_train)
bad_pred = knn_bad.predict(X_test)
bad_acc = accuracy_score(y_test, bad_pred)
print('Accuracy with k=100:', bad_acc)
Accuracy with k=100: 0.3
# Find the best k for our data
scores = []
for k in range(1, 16):
    model = KNeighborsClassifier(n_neighbors=k)
    model.fit(X_train, y_train)
    y_pred = model.predict(X_test)
    scores.append(accuracy_score(y_test, y_pred))
print('K values:', list(range(1, 16)))
print('Accuracies:', scores)
K values: [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15]
Accuracies: [1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 0.9666666666666667, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0, 1.0]
# Mini project: Let's classify unknown flowers
new_flowers = [[5, 3.2, 1.1, 0.1], [6, 2.9, 4.5, 1.5], [7, 3, 6.3, 2.5]]
results = knn.predict(new_flowers)
for i, res in enumerate(results):
    print(f"Flower {i+1} is predicted as:", iris['target_names'][res])
    
Flower 1 is predicted as: setosa
Flower 2 is predicted as: versicolor
Flower 3 is predicted as: virginica
# Best practices: Scale features for KNN
from sklearn.preprocessing import StandardScaler

scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)
X_test_scaled = scaler.transform(X_test)

knn_scaled = KNeighborsClassifier(n_neighbors=3)
knn_scaled.fit(X_train_scaled, y_train)
scaled_acc = accuracy_score(y_test, knn_scaled.predict(X_test_scaled))
print('Accuracy (with scaling):', scaled_acc)
Accuracy (with scaling): 1.0
# Troubleshooting: What if we get an error?
try:
    test_pred = knn.predict([[5, 3]])  # only two features instead of four!
except Exception as e:
    print('Oops! Error:', str(e))
    
Oops! Error: X has 2 features, but KNeighborsClassifier is expecting 4 features as input.
# Extra tip: Use kneighbors to look up the closest data points.
distances, indices = knn.kneighbors([X_test.iloc[0]])
print('Distances:', distances)
print('Indices:', indices)
Distances: [[0.2236068  0.3        0.43588989]]
Indices: [[79 90 39]]
# Challenge: Try changing the dataset or k value.
print('Try using Wine or Breast Cancer datasets with K N N.')
print('Or try new values for k, like 5 or 7!')
Try using Wine or Breast Cancer datasets with K N N.
Or try new values for k, like 5 or 7!

Recap#

We introduced KNN and saw how to use it for classification.

You learned how to:

  • Prepare data
  • Train and test a model
  • Tune k
  • Use real-world data

Good job!

Thanks for Learning KNN!#

Try KNN on your own and tell us what you discover.

If this lesson helped, like or subscribe for more Python!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.