Lesson 12 · ML algorithms deep dive
Comparing Every ML Algorithm: Capstone Project | ML Algorithms #12
The finale: every real algorithm from this series, run head to head on one final real dataset. Real wine chemistry data, three real cultivars, one fair,…
- CourseML algorithms deep dive
- Lesson12 of 12
- Video15 min
- FormatJupyter notebook · 9 code cells
- Data1 dataset
What you'll learn
Datasets used in this lesson
Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.
- wine.csv12.0 KB
📓 Full notebook
Download .ipynbML Algorithms Deep-Dive, Video 12: Capstone, Comparing Every Algorithm#
- The finale: every real algorithm from this series, run head to head on one final real dataset.
- Real wine chemistry data, three real cultivars, one fair, real comparison.
- Let's get into it.
Before You Start#
- Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
- You'll need pandas, NumPy, Matplotlib, and scikit-learn.
- Place wine.csv in the same folder as this notebook.
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
from sklearn.linear_model import LogisticRegression
from sklearn.neighbors import KNeighborsClassifier
from sklearn.tree import DecisionTreeClassifier
from sklearn.ensemble import RandomForestClassifier, GradientBoostingClassifier
from sklearn.svm import SVC
from sklearn.naive_bayes import GaussianNB
from sklearn.model_selection import train_test_split, cross_val_score, StratifiedKFold
from sklearn.preprocessing import StandardScaler
from sklearn.metrics import confusion_matrix
wine = pd.read_csv('wine.csv')
feature_cols = [c for c in wine.columns if c != 'target']
print(f'Real wines in this dataset: {len(wine)}')
print(f'Real chemical features: {len(feature_cols)}')
Part 1: A Fair Playing Field#
X = wine[feature_cols].values
y = wine['target'].values
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, random_state=42, stratify=y)
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)
X_test_scaled = scaler.transform(X_test)
skf = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
print(f'Real training wines: {len(X_train)}')
print(f'Real test wines: {len(X_test)}')
Part 2: Running Every Real Algorithm#
models = {
'Logistic Regression': (LogisticRegression(max_iter=1000), True),
'KNN (k=5)': (KNeighborsClassifier(n_neighbors=5), True),
'Decision Tree': (DecisionTreeClassifier(random_state=42), False),
'Random Forest': (RandomForestClassifier(n_estimators=100, random_state=42), False),
'Gradient Boosting': (GradientBoostingClassifier(random_state=42), False),
'SVM (RBF)': (SVC(kernel='rbf', random_state=42), True),
'Naive Bayes': (GaussianNB(), True),
}
print(f'Real algorithms entered into this comparison: {len(models)}')
results = {}
for name, (model, needs_scaling) in models.items():
X_tr = X_train_scaled if needs_scaling else X_train
X_te = X_test_scaled if needs_scaling else X_test
cv_scores = cross_val_score(model, X_tr, y_train, cv=skf, scoring='accuracy')
model.fit(X_tr, y_train)
test_accuracy = model.score(X_te, y_test)
results[name] = (cv_scores.mean(), cv_scores.std(), test_accuracy)
print(f'{name}: real CV mean={cv_scores.mean():.4f}, real CV std={cv_scores.std():.4f}, real test acc={test_accuracy:.4f}')
Part 3: Visualizing the Real Comparison#
names = list(results.keys())
means = [results[n][0] for n in names]
stds = [results[n][1] for n in names]
plt.figure(figsize=(11, 6))
plt.bar(names, means, yerr=stds, capsize=5, color='steelblue')
plt.title('Real Cross-Validated Accuracy by Algorithm')
plt.xlabel('Algorithm')
plt.ylabel('CV Accuracy')
plt.xticks(rotation=30, ha='right')
plt.ylim(0.8, 1.02)
plt.tight_layout()
plt.savefig('algorithm_comparison.png', dpi=150)
plt.show()
Part 4: Is the Real Gap at the Top Actually Meaningful?#
svm_folds = cross_val_score(SVC(kernel='rbf', random_state=42), X_train_scaled, y_train, cv=skf, scoring='accuracy')
rf_folds = cross_val_score(RandomForestClassifier(n_estimators=100, random_state=42), X_train, y_train, cv=skf, scoring='accuracy')
logreg_folds = cross_val_score(LogisticRegression(max_iter=1000), X_train_scaled, y_train, cv=skf, scoring='accuracy')
print(f'Real SVM per-fold scores: {[round(s, 4) for s in svm_folds]}')
print(f'Real Random Forest per-fold scores: {[round(s, 4) for s in rf_folds]}')
print(f'Real Logistic Regression per-fold scores: {[round(s, 4) for s in logreg_folds]}')
All three real top algorithms are essentially tied, each one trips on the same one or two genuinely hard real training wines, not on any systematic real weakness. A CV mean difference of a percentage point or two, on only 133 real training wines split five ways, is well within the kind of real noise a single hard fold can produce. This is the exact same real caution from video seven's grid search: don't over-read small real differences at the top of a leaderboard.
Part 5: Checking the Real Held-Out Test Set#
svm_final = SVC(kernel='rbf', random_state=42).fit(X_train_scaled, y_train)
svm_pred = svm_final.predict(X_test_scaled)
print('Real SVM test confusion matrix:')
print(confusion_matrix(y_test, svm_pred))
logreg_final = LogisticRegression(max_iter=1000).fit(X_train_scaled, y_train)
logreg_pred = logreg_final.predict(X_test_scaled)
print('Real logistic regression test confusion matrix:')
print(confusion_matrix(y_test, logreg_pred))
Part 6: The Real Final Recommendation#
Recommendation: logistic regression. It matched SVM and Random Forest's real accuracy, including a real perfect score on the held-out test set, while training faster and, per video two, offering real coefficients that a human can actually read and explain. SVM remains a completely legitimate real alternative if a genuinely non-linear boundary were ever needed, and Random Forest if real feature importance rankings mattered more than raw coefficients. The real, honest takeaway from all twelve videos: reach for the simplest real algorithm that gets the job done, and only add complexity once a real, validated need for it actually shows up.
Part 7: Saving Your Work#
import joblib
joblib.dump(logreg_final, 'capstone_final_model.joblib')
reloaded_model = joblib.load('capstone_final_model.joblib')
print(f'Real original model test accuracy: {logreg_final.score(X_test_scaled, y_test):.4f}')
print(f'Real reloaded model test accuracy: {reloaded_model.score(X_test_scaled, y_test):.4f}')
results_df = pd.DataFrame([
{'Algorithm': n, 'CV_Mean_Accuracy': round(v[0], 4), 'CV_Std': round(v[1], 4), 'Test_Accuracy': round(v[2], 4)}
for n, v in results.items()
])
results_df.to_csv('algorithm_comparison_results.csv', index=False)
print('Real comparison table saved to algorithm_comparison_results.csv')
print(results_df)
Wrap-Up: The Series, End to End#
- Twelve videos, one real algorithm each, tested on real cars, real tumors, real flowers, real mushrooms, real wine, real patients, real customers, and real handwritten digits.
- Regression: linear, then gradient-boosted, with an honest case where simplicity won.
- Classification: logistic regression, KNN, decision trees, random forests, SVMs, and Naive Bayes, each rematched against a real neighbor for a genuine, direct comparison.
- Unsupervised learning: K-Means and hierarchical clustering, independently converging on the same real customer segments.
- PCA, compressing real high-dimensional data without losing what actually mattered.
- This capstone, running every real classifier head to head and choosing the simplest real winner among a genuine statistical tie.
- That's the full ML Algorithms Deep-Dive series. Thanks for real building all twelve of these with me. Subscribe for what's coming next see you there.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



