Lesson 44 · Python Fundamentals
Understanding K Nearest Neighbors (KNN) in Python: A Step-by-Step Guide for Classification
In this lesson, we will learn how to use Python to find patterns and make predictions with a simple but powerful machine learning technique called KNN. You…
- CoursePython Fundamentals
- Lesson44 of 22
- Video13 min
- FormatJupyter notebook · 21 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynb
Welcome to K Nearest Neighbors (KNN) in Python#
In this lesson, we will learn how to use Python to find patterns and make predictions with a simple but powerful machine learning technique called KNN.
You will see how to work with real data, step by step.
No experience needed!
# Let's start with a quick setup.
import warnings
warnings.filterwarnings("ignore")
print("Ready to get started!")
What is KNN?#
KNN is a way for computers to recognize patterns by comparing things.
It asks: What are the most similar examples to this new one?
It is useful for classifying things (like sorting emails as spam or not spam!).
# Let's try some basic math first!
a = 9
b = 16
result = (a + b) / 2
print("Average:", result)
# Now let's use a list.
numbers = [10, 20, 30, 40, 50]
print("First number:", numbers[0])
print("Last number:", numbers[-1])
How Does KNN Work?#
KNN finds the "K" closest neighbors to any point.
It counts which category shows up the most nearby.
Then, it guesses that your new data belongs to that category.
# Let's measure distance by hand with two numbers.
x1 = 7
x2 = 17
distance = abs(x1 - x2)
print("Distance:", distance)
# Now let's use two features, like height and weight.
person_A = [170, 65]
person_B = [160, 60]
diff_height = person_A[0] - person_B[0]
diff_weight = person_A[1] - person_B[1]
euclidean_distance = (diff_height ** 2 + diff_weight ** 2) ** 0.5
print("Distance between A and B:", euclidean_distance)
Let's Work With Real Data!#
We will use the famous Iris flower dataset.
We want to predict the species of a flower from some measurements.
# Data setup
from sklearn.datasets import load_iris
import pandas as pd
iris = load_iris(as_frame=True)
df = iris['frame']
print('Shape:', df.shape)
df.head()
# Check the target distribution
df['target'].value_counts()
# Split the data into a training set and a test set
from sklearn.model_selection import train_test_split
X = df.drop('target', axis=1)
y = df['target']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
print('Training samples:', len(X_train))
print('Test samples:', len(X_test))
# Train a KNN model
from sklearn.neighbors import KNeighborsClassifier
knn = KNeighborsClassifier(n_neighbors=3)
knn.fit(X_train, y_train)
print('KNN model trained!')
# Predict on test data
y_pred = knn.predict(X_test)
print('Predicted labels:', y_pred[:10])
# Measure accuracy
from sklearn.metrics import accuracy_score
accuracy = accuracy_score(y_test, y_pred)
print("Accuracy:", accuracy)
# Try predicting a new flower using input()
values = []
feature_names = list(X.columns)
for name in feature_names:
val = float(input(f"Enter {name}: "))
values.append(val)
prediction = knn.predict([values])[0]
print("Predicted species:", iris['target_names'][prediction])
# What if we pick a bad value for k?
knn_bad = KNeighborsClassifier(n_neighbors=100)
knn_bad.fit(X_train, y_train)
bad_pred = knn_bad.predict(X_test)
bad_acc = accuracy_score(y_test, bad_pred)
print('Accuracy with k=100:', bad_acc)
# Find the best k for our data
scores = []
for k in range(1, 16):
model = KNeighborsClassifier(n_neighbors=k)
model.fit(X_train, y_train)
y_pred = model.predict(X_test)
scores.append(accuracy_score(y_test, y_pred))
print('K values:', list(range(1, 16)))
print('Accuracies:', scores)
# Mini project: Let's classify unknown flowers
new_flowers = [[5, 3.2, 1.1, 0.1], [6, 2.9, 4.5, 1.5], [7, 3, 6.3, 2.5]]
results = knn.predict(new_flowers)
for i, res in enumerate(results):
print(f"Flower {i+1} is predicted as:", iris['target_names'][res])
# Best practices: Scale features for KNN
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)
X_test_scaled = scaler.transform(X_test)
knn_scaled = KNeighborsClassifier(n_neighbors=3)
knn_scaled.fit(X_train_scaled, y_train)
scaled_acc = accuracy_score(y_test, knn_scaled.predict(X_test_scaled))
print('Accuracy (with scaling):', scaled_acc)
# Troubleshooting: What if we get an error?
try:
test_pred = knn.predict([[5, 3]]) # only two features instead of four!
except Exception as e:
print('Oops! Error:', str(e))
# Extra tip: Use kneighbors to look up the closest data points.
distances, indices = knn.kneighbors([X_test.iloc[0]])
print('Distances:', distances)
print('Indices:', indices)
# Challenge: Try changing the dataset or k value.
print('Try using Wine or Breast Cancer datasets with K N N.')
print('Or try new values for k, like 5 or 7!')
Recap#
We introduced KNN and saw how to use it for classification.
You learned how to:
- Prepare data
- Train and test a model
- Tune k
- Use real-world data
Good job!
Thanks for Learning KNN!#
Try KNN on your own and tell us what you discover.
If this lesson helped, like or subscribe for more Python!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



