Mathew K Analytics

Lesson 26 · Python For Machine Learning

Master Support Vector Machines (SVM) in Python for Advanced Machine Learning

Welcome to this hands-on Python lesson! Today, we will explore Support Vector Machines, or SVMs. SVMs are a type of machine learning model used for tasks…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb
 

Support Vector Machines (SVM) for Beginners#

Welcome to this hands-on Python lesson!

Today, we will explore Support Vector Machines, or SVMs.

SVMs are a type of machine learning model used for tasks like classifying emails or images.

We will start from the basics and build up to using SVMs on real data.

Let us get started!

# Let us start by making sure warnings will not disturb our lesson
import warnings
warnings.filterwarnings('ignore')

What is a Support Vector Machine?#

An SVM is a supervised machine learning method.

It can predict if something belongs to category A or category B.

SVM tries to find the best dividing line between different categories.

You will see how it works soon.

# First, let us import the basic libraries we will use
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
from sklearn.model_selection import train_test_split
from sklearn.svm import SVC

Let us use a Real Dataset#

We will use the Iris dataset.

It is small, famous, and lets us practice classification.

Iris data has flower measurements and a type for each flower.

# Data setup
from sklearn.datasets import load_iris
iris = load_iris(as_frame=True)
df = iris.frame
print('Shape:', df.shape)
df.head()
Shape: (150, 5)
sepal length (cm) sepal width (cm) petal length (cm) petal width (cm) target
0 5.1 3.5 1.4 0.2 0
1 4.9 3.0 1.4 0.2 0
2 4.7 3.2 1.3 0.2 0
3 4.6 3.1 1.5 0.2 0
4 5.0 3.6 1.4 0.2 0
# Let us check what we will predict
df['target'].value_counts()
target
0    50
1    50
2    50
Name: count, dtype: int64

How does SVM work?#

SVM draws lines to separate data into categories.

The line has to do the best job splitting groups apart.

SVM can also work in more than two dimensions.

Let us try a tiny SVM example to see how training works.

# Training our first SVM using only two features for visualization
X = df[['sepal length (cm)', 'sepal width (cm)']]
y = df['target']
svm_model = SVC(kernel='linear')
svm_model.fit(X, y)
SVC(kernel='linear')
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
# Visualize data points and the decision boundary
plt.figure(figsize=(8,6))
plt.scatter(X['sepal length (cm)'], X['sepal width (cm)'], c=y, cmap='viridis', s=40, edgecolors='k')
plt.xlabel('Sepal Length (cm)')
plt.ylabel('Sepal Width (cm)')
plt.title('Iris Data and SVM Decision Boundary')

# Plotting the decision boundary (only works for linear SVM and 2 features)
w = svm_model.coef_[0]
b = svm_model.intercept_[0]
x_plot = np.linspace(X['sepal length (cm)'].min(), X['sepal length (cm)'].max(), 30)
y_plot = -(w[0]/w[1]) * x_plot - b/w[1]
plt.plot(x_plot, y_plot, 'r--')
plt.show()
No description has been provided for this image
# Let us split the data into training and testing sets
X = df.drop('target', axis=1)
y = df['target']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)
print('Training size:', X_train.shape)
print('Testing size:', X_test.shape)
Training size: (105, 4)
Testing size: (45, 4)
# Now, train an SVM on the train data
classifier = SVC(kernel='linear', random_state=42)
classifier.fit(X_train, y_train)
SVC(kernel='linear', random_state=42)
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
# Let us test how well the model predicts on the test data
predictions = classifier.predict(X_test)
print('First 10 predicted:', predictions[:10].tolist())
print('Actual:', y_test.values[:10].tolist())
First 10 predicted: [1, 0, 2, 1, 1, 0, 1, 2, 1, 1]
Actual: [1, 0, 2, 1, 1, 0, 1, 2, 1, 1]
# Let us measure model accuracy
from sklearn.metrics import accuracy_score
accuracy = accuracy_score(y_test, predictions)
print('Accuracy:', accuracy)
Accuracy: 1.0
# Try SVM with a non-linear kernel
nonlinear_svm = SVC(kernel='rbf', random_state=42)
nonlinear_svm.fit(X_train, y_train)
rbf_predictions = nonlinear_svm.predict(X_test)
rbf_accuracy = accuracy_score(y_test, rbf_predictions)
print('RBF Kernel Accuracy:', rbf_accuracy)
RBF Kernel Accuracy: 1.0
# What happens if we change the C parameter? Try asking the user!
c_value = float(input('Pick a value for C (try 0.1, 1, or 10): '))
c_svm = SVC(kernel='linear', C=c_value, random_state=42)
c_svm.fit(X_train, y_train)
c_predictions = c_svm.predict(X_test)
c_acc = accuracy_score(y_test, c_predictions)
print('Accuracy with C =', c_value, ':', c_acc)
Accuracy with C = 1.0 : 1.0
 
# Using SVM with only two flower classes
binary_df = df[df['target'].isin([0, 1])]
X_bin = binary_df.drop('target', axis=1)
y_bin = binary_df['target']
X_train_bin, X_test_bin, y_train_bin, y_test_bin = train_test_split(X_bin, y_bin, test_size=0.3, random_state=40)
binary_svm = SVC(kernel='linear', random_state=40)
binary_svm.fit(X_train_bin, y_train_bin)
bin_preds = binary_svm.predict(X_test_bin)
bin_acc = accuracy_score(y_test_bin, bin_preds)
print('Binary Classification Accuracy:', bin_acc)
Binary Classification Accuracy: 1.0
# Let us check which predictions are wrong
incorrect = np.where(predictions != y_test)[0]
print('Indices of mistakes:', incorrect.tolist())
print('Predicted:', predictions[incorrect])
print('Actual:', y_test.values[incorrect])
Indices of mistakes: []
Predicted: []
Actual: []
# Let us use SVM to predict a new flower
print('Please enter four numbers for sepal length, sepal width, petal length, and petal width (separated by spaces): ')
user_input = input()
new_data = np.array([float(x) for x in user_input.strip().split()]).reshape(1, -1)
prediction = classifier.predict(new_data)[0]
print('The SVM predicts this flower is type:', iris.target_names[prediction])
Please enter four numbers for sepal length, sepal width, petal length, and petal width (separated by spaces): 
The SVM predicts this flower is type: setosa
 
# Practice: Write a loop that prints all SVM model support vectors
for vec in classifier.support_vectors_:
    print(vec)
    
[4.8 3.4 1.9 0.2]
[5.1 3.3 1.7 0.5]
[4.5 2.3 1.3 0.3]
[5.6 3.  4.5 1.5]
[5.4 3.  4.5 1.5]
[6.7 3.  5.  1.7]
[5.9 3.2 4.8 1.8]
[5.1 2.5 3.  1.1]
[6.  2.7 5.1 1.6]
[6.3 2.5 4.9 1.5]
[6.1 2.9 4.7 1.4]
[6.5 2.8 4.6 1.5]
[6.9 3.1 4.9 1.5]
[6.3 2.3 4.4 1.3]
[6.3 2.8 5.1 1.5]
[6.3 2.7 4.9 1.8]
[6.  3.  4.8 1.8]
[6.  2.2 5.  1.5]
[6.2 2.8 4.8 1.8]
[6.5 3.  5.2 2. ]
[7.2 3.  5.8 1.6]
[5.6 2.8 4.9 2. ]
[5.9 3.  5.1 1.8]
[4.9 2.5 4.5 1.7]

Recap: What Have We Learned?#

  • What SVMs are and why we use them.
  • How to load data and fit an SVM model.
  • How to measure model performance.
  • That SVMs work for both two and more classes.
  • Changing parameters can change model results.

Great job making it to the end!

Challenge: Try Your Own!#

Pick new test_size and kernel values.

Use another dataset from sklearn (like wine or digits) with an SVM.

Share your results and what you learned in the comments below.

If you liked this lesson, subscribe for more coding videos!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.