Lesson 2 · ML algorithms deep dive
Logistic Regression for Classification | ML Algorithms #2
Video two of the 12-part series: the first real classification algorithm. Real diagnostic measurements from the Wisconsin breast cancer dataset, a genuine,…
- CourseML algorithms deep dive
- Lesson2 of 12
- Video21 min
- FormatJupyter notebook · 14 code cells
- Data1 dataset
What you'll learn
Datasets used in this lesson
Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.
- breast_cancer.csv118.5 KB
📓 Full notebook
Download .ipynbML Algorithms Deep-Dive, Video 2: Logistic Regression#
- Video two of the 12-part series: the first real classification algorithm.
- Real diagnostic measurements from the Wisconsin breast cancer dataset, a genuine, widely-used medical dataset.
- Let's get into it.
Before You Start#
- Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
- You'll need pandas, NumPy, Matplotlib, and scikit-learn.
- Place breast_cancer.csv in the same folder as this notebook.
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split, cross_val_score, StratifiedKFold
from sklearn.preprocessing import StandardScaler
from sklearn.metrics import confusion_matrix, precision_score, recall_score, f1_score, roc_auc_score, roc_curve, accuracy_score
cancer = pd.read_csv('breast_cancer.csv')
print(f'Real patients in this dataset: {len(cancer)}')
print(cancer['target'].value_counts())
Part 1: Why Not Just Use Linear Regression?#
cancer['malignant'] = (cancer['target'] == 0).astype(int)
print(cancer['malignant'].value_counts())
print(f"Real malignant rate: {cancer['malignant'].mean():.4f}")
Part 2: The Sigmoid Function#
z_values = np.array([-2, -1, 0, 1, 2])
sigmoid_values = 1 / (1 + np.exp(-z_values))
for z, p in zip(z_values, sigmoid_values):
print(f'z={z}: sigmoid(z)={p:.4f}')
plt.figure(figsize=(8, 5))
z_range = np.linspace(-6, 6, 200)
plt.plot(z_range, 1 / (1 + np.exp(-z_range)), color='steelblue')
plt.axhline(0.5, color='gray', linestyle='--', linewidth=1)
plt.axvline(0, color='gray', linestyle='--', linewidth=1)
plt.title('The Sigmoid Function')
plt.xlabel('z (the linear combination of features)')
plt.ylabel('Predicted Probability')
plt.show()
Part 3: Fitting a Real Model#
features = ['mean radius', 'mean texture', 'mean concavity', 'mean smoothness']
X = cancer[features].values
y = cancer['malignant'].values
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, random_state=42, stratify=y)
print(f'Real training patients: {len(X_train)}')
print(f'Real test patients: {len(X_test)}')
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)
X_test_scaled = scaler.transform(X_test)
model = LogisticRegression(max_iter=1000).fit(X_train_scaled, y_train)
print(f'Real intercept: {model.intercept_[0]:.4f}')
for name, coef in zip(features, model.coef_[0]):
print(f'Real coefficient for {name}: {coef:.4f}')
Each real coefficient here is a change in log-odds, not a change in probability directly; a coefficient of 3.42 for radius means each one real standard deviation increase in radius multiplies the real odds of malignancy by about e to the 3.42, roughly 30 times. That's a real, large effect, consistent with tumor size being one of the strongest real predictors of malignancy.
Part 4: The Confusion Matrix#
y_pred = model.predict(X_test_scaled)
cm = confusion_matrix(y_test, y_pred)
print('Real confusion matrix (rows=actual benign/malignant, cols=predicted benign/malignant):')
print(cm)
true_neg, false_pos, false_neg, true_pos = cm.ravel()
print(f'Real true negatives (correctly benign): {true_neg}')
print(f'Real false positives (benign called malignant): {false_pos}')
print(f'Real false negatives (malignant missed): {false_neg}')
print(f'Real true positives (correctly caught malignant): {true_pos}')
Part 5: Precision, Recall, and F1#
accuracy = accuracy_score(y_test, y_pred)
precision = precision_score(y_test, y_pred)
recall = recall_score(y_test, y_pred)
f1 = f1_score(y_test, y_pred)
print(f'Real accuracy: {accuracy:.4f}')
print(f'Real precision: {precision:.4f}')
print(f'Real recall: {recall:.4f}')
print(f'Real F1 score: {f1:.4f}')
In a real medical screening context, recall usually matters more than precision: a missed real cancer is far more costly than a false alarm that a follow-up test can rule out. This real model's default recall of 0.89 means about 11 percent of real malignant cases would be missed, which is exactly the kind of number worth trying to improve before precision.
Part 6: The Threshold Tradeoff#
y_proba = model.predict_proba(X_test_scaled)[:, 1]
for threshold in [0.5, 0.3, 0.2]:
y_pred_t = (y_proba >= threshold).astype(int)
p = precision_score(y_test, y_pred_t)
r = recall_score(y_test, y_pred_t)
cm_t = confusion_matrix(y_test, y_pred_t)
print(f'threshold={threshold}: real precision={p:.4f}, real recall={r:.4f}, real false negatives={cm_t[1, 0]}, real false positives={cm_t[0, 1]}')
Part 7: The ROC Curve and AUC#
fpr, tpr, thresholds = roc_curve(y_test, y_proba)
auc = roc_auc_score(y_test, y_proba)
print(f'Real AUC: {auc:.4f}')
plt.figure(figsize=(6, 6))
plt.plot(fpr, tpr, color='steelblue', label=f'Real model (AUC={auc:.3f})')
plt.plot([0, 1], [0, 1], color='gray', linestyle='--', label='Random guessing')
plt.title('Real ROC Curve')
plt.xlabel('False Positive Rate')
plt.ylabel('True Positive Rate')
plt.legend()
plt.show()
Part 8: Scaling Up to All 30 Real Features#
feature_cols = [c for c in cancer.columns if c not in ('target', 'malignant')]
X_full = cancer[feature_cols].values
X_train_full, X_test_full, y_train_full, y_test_full = train_test_split(X_full, y, test_size=0.25, random_state=42, stratify=y)
scaler_full = StandardScaler()
X_train_full_scaled = scaler_full.fit_transform(X_train_full)
X_test_full_scaled = scaler_full.transform(X_test_full)
full_model = LogisticRegression(max_iter=5000).fit(X_train_full_scaled, y_train_full)
full_auc = roc_auc_score(y_test_full, full_model.predict_proba(X_test_full_scaled)[:, 1])
print(f'Real 4-feature AUC: {auc:.4f}')
print(f'Real 30-feature AUC: {full_auc:.4f}')
print(f'Real 30-feature accuracy: {full_model.score(X_test_full_scaled, y_test_full):.4f}')
Part 9: Regularization Strength#
skf = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
for C in [0.01, 0.1, 1.0, 10.0, 100.0]:
scores = cross_val_score(LogisticRegression(max_iter=5000, C=C), X_train_full_scaled, y_train_full, cv=skf, scoring='roc_auc')
print(f'C={C}: real mean CV AUC={scores.mean():.4f}')
More real features didn't call for less regularization, it called for more. C equals 100 gave the model the most real freedom to fit the training data closely, and it generalized the worst of the five real settings tested. That's the same real theme from video one: the setting that fits training data best is not automatically the setting that predicts new real data best.
Part 10: Strengths, Weaknesses, and When to Use It#
- Strength: outputs a genuine, calibrated real probability, not just a hard label.
- Strength: coefficients are interpretable as real effects on log-odds.
- Strength: the decision threshold can be tuned after training to match real-world costs, as part six showed.
- Weakness: assumes a real linear decision boundary in log-odds space; genuinely non-linear real patterns need a different algorithm.
- Use it as a real, fast, interpretable baseline for any binary classification problem before reaching for something more complex.
Part 11: Saving Your Work#
final_model = LogisticRegression(max_iter=5000, C=0.1).fit(X_train_full_scaled, y_train_full)
final_proba = final_model.predict_proba(X_test_full_scaled)[:, 1]
final_auc = roc_auc_score(y_test_full, final_proba)
final_fpr, final_tpr, _ = roc_curve(y_test_full, final_proba)
plt.figure(figsize=(6, 6))
plt.plot(final_fpr, final_tpr, color='steelblue', label=f'Final model (AUC={final_auc:.3f})')
plt.plot([0, 1], [0, 1], color='gray', linestyle='--')
plt.title('Real Final Model ROC Curve')
plt.xlabel('False Positive Rate')
plt.ylabel('True Positive Rate')
plt.legend()
plt.savefig('logistic_regression_final_roc.png', dpi=150)
plt.show()
import joblib
joblib.dump(final_model, 'logistic_regression_model.joblib')
reloaded_model = joblib.load('logistic_regression_model.joblib')
print(f'Real original model test accuracy: {final_model.score(X_test_full_scaled, y_test_full):.4f}')
print(f'Real reloaded model test accuracy: {reloaded_model.score(X_test_full_scaled, y_test_full):.4f}')
Wrap-Up: What You Learned#
- The real sigmoid function, and why it's the right tool for turning a linear score into a real probability.
- Fitting real logistic regression and interpreting its coefficients as real log-odds effects.
- The confusion matrix, and the real difference between precision, recall, and F1.
- Tuning the real decision threshold to trade precision for recall, a genuinely practical, real-world lever.
- The ROC curve and AUC as a threshold-free real summary of model quality.
- Regularization strength, and a real case where the least-constrained model generalized worst.
- Video three moves to K-Nearest Neighbors, a completely different kind of classifier, on the real Iris dataset. Subscribe so it lands automatically see you there.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



