Mathew K Analytics

Lesson 32 · Data Mining

Support Vector Machines (SVM): Principles, Methods, and Practical Applications

Welcome to our beginner-friendly lesson on Support Vector Machines, also known as SVM! We will use the Pima Indians Diabetes Dataset to walk through real…

⬇ Download notebookOpen in Colab ↗

📓 Full notebook

Download .ipynb

Week 89: Introduction to Support Vector Machines (SVM)#

Welcome to our beginner-friendly lesson on Support Vector Machines, also known as SVM!

We will use the Pima Indians Diabetes Dataset to walk through real data mining workflows.

By the end, you will know how to apply SVM for classification, tune it, interpret results, and spot common issues.

Let us get started!

# Suppress warnings for a cleaner learning environment
import warnings; warnings.filterwarnings('ignore')
import numpy as np
np.random.seed(42)

What are Support Vector Machines (SVM)?#

Support Vector Machines are supervised learning methods used mostly for classification tasks.

They work by finding the best boundary that separates different categories in your data.

SVM is powerful for handling data that is not easily separated, thanks to the ability to use kernels.

Real-world uses include medical diagnosis, face recognition, and text sorting, just like classifying diabetes status here.

# Data setup (Pima Indians Diabetes Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/jbrownlee/Datasets/master/pima-indians-diabetes.data.csv'
cols = ['Pregnancies','Glucose','BloodPressure','SkinThickness','Insulin','BMI','DiabetesPedigree','Age','Outcome']
df = pd.read_csv(url, names=cols)
print(df.shape)
print(df.head(3))
(768, 9)
   Pregnancies  Glucose  BloodPressure  SkinThickness  Insulin   BMI  \
0            6      148             72             35        0  33.6   
1            1       85             66             29        0  26.6   
2            8      183             64              0        0  23.3   

   DiabetesPedigree  Age  Outcome  
0             0.627   50        1  
1             0.351   31        0  
2             0.672   32        1  
# Quick check for missing data or unusual values
print(df.isnull().sum())
print('Zeroes in columns:')
print((df == 0).sum())
Pregnancies         0
Glucose             0
BloodPressure       0
SkinThickness       0
Insulin             0
BMI                 0
DiabetesPedigree    0
Age                 0
Outcome             0
dtype: int64
Zeroes in columns:
Pregnancies         111
Glucose               5
BloodPressure        35
SkinThickness       227
Insulin             374
BMI                  11
DiabetesPedigree      0
Age                   0
Outcome             500
dtype: int64
# Replace zero values in specific columns with the median (except 'Pregnancies' and 'Outcome')
cols_with_zero_na = ['Glucose','BloodPressure','SkinThickness','Insulin','BMI']
for col in cols_with_zero_na:
    df[col] = df[col].replace(0, df[col].median())
print(df[cols_with_zero_na].head(3))
   Glucose  BloodPressure  SkinThickness  Insulin   BMI
0      148             72             35     30.5  33.6
1       85             66             29     30.5  26.6
2      183             64             23     30.5  23.3
# Splitting data into input features and target label
X = df.drop('Outcome', axis=1)
y = df['Outcome']
print(X.shape, y.shape)
(768, 8) (768,)

Why Do We Split Data into Training and Testing?#

SVM, like most models, needs to be evaluated fairly.

We use some data to train (learn patterns), and separate, unseen data to test (check real-world accuracy).

This avoids overfitting, when a model is only good at remembering, not generalizing.

# Split data: 75% training, 25% testing
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, random_state=42)
print(X_train.shape, X_test.shape)
(576, 8) (192, 8)
# Standardize features for SVM
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)
X_test_scaled = scaler.transform(X_test)

Fitting a Simple SVM Classifier#

Let us train our first SVM model!

We will use scikit-learn's SVC, which stands for Support Vector Classifier.

The default kernel is 'rbf', which helps handle non-linear splits.

# Train an SVM model
from sklearn.svm import SVC
svm_model = SVC(random_state=42)
svm_model.fit(X_train_scaled, y_train)
SVC(random_state=42)
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
# Make predictions and evaluate accuracy
from sklearn.metrics import accuracy_score
y_pred = svm_model.predict(X_test_scaled)
acc = accuracy_score(y_test, y_pred)
print(f'SVM accuracy: {acc:.2%}')
SVM accuracy: 74.48%
# View confusion matrix for deeper insights
from sklearn.metrics import confusion_matrix
cm = confusion_matrix(y_test, y_pred)
print(cm)
[[102  21]
 [ 28  41]]

SVM Kernels Exploring Linear Versus RBF#

Kernels let SVM find boundaries that are not just straight lines.

A linear kernel works when points can be separated by a line. An 'rbf' (radial basis function) kernel handles curved boundaries.

Let us compare both to see the impact.

# Try linear kernel SVM
svm_linear = SVC(kernel='linear', random_state=42)
svm_linear.fit(X_train_scaled, y_train)
y_pred_linear = svm_linear.predict(X_test_scaled)
print('Linear kernel accuracy:', accuracy_score(y_test, y_pred_linear))
Linear kernel accuracy: 0.734375
# Tune the penalty parameter C to see its effect
svm_c_high = SVC(C=5.0, random_state=42)
svm_c_high.fit(X_train_scaled, y_train)
acc_high = accuracy_score(y_test, svm_c_high.predict(X_test_scaled))
svm_c_low = SVC(C=0.1, random_state=42)
svm_c_low.fit(X_train_scaled, y_train)
acc_low = accuracy_score(y_test, svm_c_low.predict(X_test_scaled))
print(f'High C accuracy: {acc_high:.2%}')
print(f'Low C accuracy: {acc_low:.2%}')
High C accuracy: 67.19%
Low C accuracy: 74.48%
# Try a mini grid search for hyperparameter tuning
from sklearn.model_selection import GridSearchCV
params = {'C': [0.1, 1, 10], 'kernel': ['linear', 'rbf']}
grid = GridSearchCV(SVC(random_state=42), params, cv=3)
grid.fit(X_train_scaled, y_train)
print('Best params:', grid.best_params_)
print('Best cross-validation score:', grid.best_score_)
Best params: {'C': 0.1, 'kernel': 'linear'}
Best cross-validation score: 0.7708333333333334
# ROC curve for SVM performance
from sklearn.metrics import roc_curve, auc
import matplotlib.pyplot as plt
y_scores = svm_model.decision_function(X_test_scaled)
fpr, tpr, thresholds = roc_curve(y_test, y_scores)
roc_auc = auc(fpr, tpr)
plt.plot(fpr, tpr, label=f'AUC = {roc_auc:.2f}')
plt.plot([0, 1], [0, 1], linestyle='--', color='gray')
plt.xlabel('False Positive Rate')
plt.ylabel('True Positive Rate')
plt.title('SVM ROC Curve')
plt.legend()
plt.show()
No description has been provided for this image

Real-World SVM Example: Predicting Diabetes Status#

Now that you have seen the basics, let us use what we learned in a mini-project.

Suppose a clinic wants to screen new patients quickly for diabetes risk.

We will use SVM for this classification problem.

# Simulate screening a new patient
user_input = [6, 140, 70, 28, 0, 35.0, 0.537, 41]
import numpy as np
user_input = np.array(user_input).reshape(1, -1)
user_input[:, 1:6] = np.where(user_input[:, 1:6]==0, [df[c].median() for c in cols_with_zero_na], user_input[:, 1:6])
user_input_scaled = scaler.transform(user_input)
pred = svm_model.predict(user_input_scaled)
print('Diabetes risk:' , 'Positive' if pred[0]==1 else 'Negative')
Diabetes risk: Positive

SVM Best Practices and Troubleshooting#

Always scale your features before training SVMs, especially if your data includes physical measurements.

Try out kernel and penalty options; defaults may not be best for your data.

If your SVM is slow with big datasets, try a linear SVM or other algorithms.

Pay attention to your proportions of classesSVMs can be sensitive if one group is much bigger than the other.

Check the confusion matrix and ROC curve, not just accuracy.

Tips for Learning SVMs#

Play with kernel typesand do not be afraid to try 'poly' or 'sigmoid' for variety.

If your accuracy is low, check your data cleaning and class balance.

Use tools like GridSearchCV to find good settings, but rely on test data for real-world evaluation.

Ask for helpeveryone finds tuning SVMs tricky at first!

# Self-check: Try classifying your own row! Enter numbers and run live prediction.
input_row = input('Enter 8 numbers (comma-separated) for a new patient: ')
try:
    user_vals = [float(n.strip()) for n in input_row.split(',')]
    arr = np.array(user_vals).reshape(1, -1)
    arr[:, 1:6] = np.where(arr[:, 1:6]==0, [df[c].median() for c in cols_with_zero_na], arr[:, 1:6])
    arr_scaled = scaler.transform(arr)
    result = svm_model.predict(arr_scaled)
    print('This patient is', 'Diabetes Positive' if result[0]==1 else 'Diabetes Negative')
except:
    print('Could not parse your input. Please try again.')
    
This patient is Diabetes Positive

Challenge: How Would SVM Handle Different Problems?#

Suppose you received images, text, or even sensor data instead.

What steps would you take to adapt SVM for those?

Hint: Consider feature engineering and kernel options.

Try out your answer below:

SVM Recap: What Have You Learned?#

  • You can load, clean, and split real health datasets.
  • SVMs work by finding boundaries to separate classes.
  • Feature scaling and kernel choices are key for performance.
  • Always check more than just accuracylook at confusion matrices and ROC curves.

You can now build and tune your own SVM for classification problems!

Thanks for learning SVM with us!#

Did this notebook help you? Hit like and subscribe for more hands-on machine learning lessons.

Feel free to post your SVM challenge answers and questions in the commentswe respond to all learners!

See you in the next one!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.