Mathew K Analytics

Lesson 7 · ML algorithms deep dive

Support Vector Machines Made Intuitive | ML Algorithms #7

Video seven of the 12-part series: revisiting real breast cancer data for a direct rematch against video two. Real diagnostic measurements, same real…

What you'll learn

Datasets used in this lesson

Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.

📓 Full notebook

Download .ipynb

ML Algorithms Deep-Dive, Video 7: Support Vector Machines#

  • Video seven of the 12-part series: revisiting real breast cancer data for a direct rematch against video two.
  • Real diagnostic measurements, same real patients, same real target.
  • Let's get into it.

Before You Start#

  • Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
  • You'll need pandas, NumPy, Matplotlib, and scikit-learn.
  • Place breast_cancer.csv in the same folder as this notebook.
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
from sklearn.svm import SVC
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split, cross_val_score, StratifiedKFold
from sklearn.preprocessing import StandardScaler

cancer = pd.read_csv('breast_cancer.csv')
cancer['malignant'] = (cancer['target'] == 0).astype(int)
features = ['mean radius', 'mean texture', 'mean concavity', 'mean smoothness']
print(f'Real patients: {len(cancer)}')
print(f'Real malignant rate: {cancer["malignant"].mean():.4f}')
Real patients: 569
Real malignant rate: 0.3726

Part 1: A Different Kind of Boundary#

X = cancer[features].values
y = cancer['malignant'].values
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, random_state=42, stratify=y)
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)
X_test_scaled = scaler.transform(X_test)
print(f'Real training patients: {len(X_train)}')
print(f'Real test patients: {len(X_test)}')
Real training patients: 426
Real test patients: 143

Part 2: Fitting a Real Linear SVM#

logreg = LogisticRegression(max_iter=1000).fit(X_train_scaled, y_train)
svm_linear = SVC(kernel='linear', random_state=42).fit(X_train_scaled, y_train)
print(f'Real logistic regression test accuracy: {logreg.score(X_test_scaled, y_test):.4f}')
print(f'Real linear SVM test accuracy: {svm_linear.score(X_test_scaled, y_test):.4f}')
Real logistic regression test accuracy: 0.9301
Real linear SVM test accuracy: 0.9441

Part 3: Support Vectors#

n_support = svm_linear.support_vectors_.shape[0]
print(f'Real support vectors: {n_support} of {len(X_train_scaled)} training patients')
print(f'Real support vectors per class: {svm_linear.n_support_}')
Real support vectors: 71 of 426 training patients
Real support vectors per class: [36 35]

Part 4: The Kernel Trick#

for kernel in ['linear', 'rbf', 'poly']:
    svm_kernel = SVC(kernel=kernel, random_state=42).fit(X_train_scaled, y_train)
    print(f'kernel={kernel}: real test accuracy={svm_kernel.score(X_test_scaled, y_test):.4f}')
kernel=linear: real test accuracy=0.9441
kernel=rbf: real test accuracy=0.9650
kernel=poly: real test accuracy=0.9161

Part 5: C: The Margin-Versus-Errors Tradeoff#

skf = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
for C in [0.01, 0.1, 1, 10, 100]:
    scores = cross_val_score(SVC(kernel='linear', C=C, random_state=42), X_train_scaled, y_train, cv=skf, scoring='accuracy')
    print(f'C={C}: real mean CV accuracy={scores.mean():.4f}')
C=0.01: real mean CV accuracy=0.9179
C=0.1: real mean CV accuracy=0.9366
C=1: real mean CV accuracy=0.9413
C=10: real mean CV accuracy=0.9389
C=100: real mean CV accuracy=0.9366

Part 6: Gamma: How Far Each Point's Influence Reaches#

for C in [0.1, 1, 10, 100]:
    for gamma in [0.001, 0.01, 0.1, 1]:
        scores = cross_val_score(SVC(kernel='rbf', C=C, gamma=gamma, random_state=42), X_train_scaled, y_train, cv=skf, scoring='accuracy')
        print(f'C={C}, gamma={gamma}: real mean CV accuracy={scores.mean():.4f}')
C=0.1, gamma=0.001: real mean CV accuracy=0.6268
C=0.1, gamma=0.01: real mean CV accuracy=0.7865
C=0.1, gamma=0.1: real mean CV accuracy=0.9296
C=0.1, gamma=1: real mean CV accuracy=0.9036
C=1, gamma=0.001: real mean CV accuracy=0.7888
C=1, gamma=0.01: real mean CV accuracy=0.9225
C=1, gamma=0.1: real mean CV accuracy=0.9413
C=1, gamma=1: real mean CV accuracy=0.9436
C=10, gamma=0.001: real mean CV accuracy=0.9202
C=10, gamma=0.01: real mean CV accuracy=0.9436
C=10, gamma=0.1: real mean CV accuracy=0.9319
C=10, gamma=1: real mean CV accuracy=0.9483
C=100, gamma=0.001: real mean CV accuracy=0.9389
C=100, gamma=0.01: real mean CV accuracy=0.9366
C=100, gamma=0.1: real mean CV accuracy=0.9319
C=100, gamma=1: real mean CV accuracy=0.9225

Part 7: A Real Caution About Trusting the Grid#

grid_best = SVC(kernel='rbf', C=10, gamma=1, random_state=42).fit(X_train_scaled, y_train)
print(f'Real grid-search winner test accuracy: {grid_best.score(X_test_scaled, y_test):.4f}')
default_rbf = SVC(kernel='rbf', random_state=42).fit(X_train_scaled, y_train)
print(f'Real scikit-learn default RBF test accuracy: {default_rbf.score(X_test_scaled, y_test):.4f}')
print(f'Real default gamma actually used: {default_rbf._gamma:.4f}')
Real grid-search winner test accuracy: 0.9161
Real scikit-learn default RBF test accuracy: 0.9650
Real default gamma actually used: 0.2500

This is a genuinely important, honest caveat: with only 426 real training patients split five ways, each real cross-validation fold has just roughly 85 patients, and a search across sixteen real combinations can land on a setting that happened to fit those particular real folds well without genuinely generalizing better. Grid search is a real, useful tool, not a real guarantee; it's still worth sanity-checking the winner against the true real held-out set before trusting it.

Part 8: Visualizing the Real Decision Boundary#

X2 = cancer[['mean radius', 'mean concavity']].values
X2_train, X2_test, y2_train, y2_test = train_test_split(X2, y, test_size=0.25, random_state=42, stratify=y)
scaler_2d = StandardScaler()
X2_train_scaled = scaler_2d.fit_transform(X2_train)
svm_2d = SVC(kernel='rbf', random_state=42).fit(X2_train_scaled, y2_train)
x_min, x_max = X2_train_scaled[:, 0].min() - 1, X2_train_scaled[:, 0].max() + 1
y_min, y_max = X2_train_scaled[:, 1].min() - 1, X2_train_scaled[:, 1].max() + 1
xx, yy = np.meshgrid(np.linspace(x_min, x_max, 300), np.linspace(y_min, y_max, 300))
zz = svm_2d.decision_function(np.c_[xx.ravel(), yy.ravel()]).reshape(xx.shape)
plt.figure(figsize=(8, 6))
plt.contourf(xx, yy, zz, levels=[-10, 0, 10], colors=['#fbdcdc', '#dcecfb'], alpha=0.6)
plt.contour(xx, yy, zz, levels=[-1, 0, 1], colors='black', linestyles=['--', '-', '--'])
plt.scatter(X2_train_scaled[:, 0], X2_train_scaled[:, 1], c=y2_train, cmap='coolwarm', edgecolor='black', s=25)
plt.scatter(svm_2d.support_vectors_[:, 0], svm_2d.support_vectors_[:, 1], facecolors='none', edgecolors='black', s=100, linewidths=1.5, label='Support vectors')
plt.title('Real RBF Decision Boundary, Margin, and Support Vectors')
plt.xlabel('Mean Radius (scaled)')
plt.ylabel('Mean Concavity (scaled)')
plt.legend()
plt.savefig('svm_decision_boundary.png', dpi=150)
plt.show()
No description has been provided for this image

Part 9: Strengths, Weaknesses, and When to Use It#

  • Strength: often genuinely strong on real, moderate-sized, cleanly-separated real datasets like this one.
  • Strength: kernels let it capture real curved boundaries without manual feature engineering.
  • Weakness: doesn't output a real probability by default, unlike logistic regression's natural real confidence score.
  • Weakness: two real interacting hyperparameters, C and gamma, that need genuine tuning and honest validation, as part seven showed directly.
  • Use it as a genuinely strong real alternative to logistic regression when the real decision boundary might not be a straight line.

Part 10: Saving Your Work#

import joblib
joblib.dump(default_rbf, 'svm_model.joblib')
reloaded_svm = joblib.load('svm_model.joblib')
print(f'Real original model test accuracy: {default_rbf.score(X_test_scaled, y_test):.4f}')
print(f'Real reloaded model test accuracy: {reloaded_svm.score(X_test_scaled, y_test):.4f}')
Real original model test accuracy: 0.9650
Real reloaded model test accuracy: 0.9650

Wrap-Up: What You Learned#

  • Maximum margin boundaries, and a real rematch against video two's logistic regression on identical data.
  • Support vectors: the small real subset of training points that actually define the boundary.
  • The kernel trick, and a real, honest case where the fancier polynomial kernel underperformed a simpler one.
  • C and gamma, tuned independently and then together, with the real interaction between them made visible.
  • A real, humbling result where the cross-validation winner underperformed the untuned default on the real test set.
  • Visualizing the real margin and support vectors directly on a two-feature decision boundary.
  • Video eight moves to Naive Bayes, revisiting real mushroom data for another direct algorithm rematch. Subscribe so it lands automatically see you there.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.