Mathew K Analytics

Lesson 2 · ML algorithms deep dive

Logistic Regression for Classification | ML Algorithms #2

Video two of the 12-part series: the first real classification algorithm. Real diagnostic measurements from the Wisconsin breast cancer dataset, a genuine,…

What you'll learn

Datasets used in this lesson

Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.

📓 Full notebook

Download .ipynb

ML Algorithms Deep-Dive, Video 2: Logistic Regression#

  • Video two of the 12-part series: the first real classification algorithm.
  • Real diagnostic measurements from the Wisconsin breast cancer dataset, a genuine, widely-used medical dataset.
  • Let's get into it.

Before You Start#

  • Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
  • You'll need pandas, NumPy, Matplotlib, and scikit-learn.
  • Place breast_cancer.csv in the same folder as this notebook.
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split, cross_val_score, StratifiedKFold
from sklearn.preprocessing import StandardScaler
from sklearn.metrics import confusion_matrix, precision_score, recall_score, f1_score, roc_auc_score, roc_curve, accuracy_score

cancer = pd.read_csv('breast_cancer.csv')
print(f'Real patients in this dataset: {len(cancer)}')
print(cancer['target'].value_counts())
Real patients in this dataset: 569
target
1    357
0    212
Name: count, dtype: int64

Part 1: Why Not Just Use Linear Regression?#

cancer['malignant'] = (cancer['target'] == 0).astype(int)
print(cancer['malignant'].value_counts())
print(f"Real malignant rate: {cancer['malignant'].mean():.4f}")
malignant
0    357
1    212
Name: count, dtype: int64
Real malignant rate: 0.3726

Part 2: The Sigmoid Function#

z_values = np.array([-2, -1, 0, 1, 2])
sigmoid_values = 1 / (1 + np.exp(-z_values))
for z, p in zip(z_values, sigmoid_values):
    print(f'z={z}: sigmoid(z)={p:.4f}')
z=-2: sigmoid(z)=0.1192
z=-1: sigmoid(z)=0.2689
z=0: sigmoid(z)=0.5000
z=1: sigmoid(z)=0.7311
z=2: sigmoid(z)=0.8808
plt.figure(figsize=(8, 5))
z_range = np.linspace(-6, 6, 200)
plt.plot(z_range, 1 / (1 + np.exp(-z_range)), color='steelblue')
plt.axhline(0.5, color='gray', linestyle='--', linewidth=1)
plt.axvline(0, color='gray', linestyle='--', linewidth=1)
plt.title('The Sigmoid Function')
plt.xlabel('z (the linear combination of features)')
plt.ylabel('Predicted Probability')
plt.show()
No description has been provided for this image

Part 3: Fitting a Real Model#

features = ['mean radius', 'mean texture', 'mean concavity', 'mean smoothness']
X = cancer[features].values
y = cancer['malignant'].values
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, random_state=42, stratify=y)
print(f'Real training patients: {len(X_train)}')
print(f'Real test patients: {len(X_test)}')
Real training patients: 426
Real test patients: 143
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)
X_test_scaled = scaler.transform(X_test)
model = LogisticRegression(max_iter=1000).fit(X_train_scaled, y_train)
print(f'Real intercept: {model.intercept_[0]:.4f}')
for name, coef in zip(features, model.coef_[0]):
    print(f'Real coefficient for {name}: {coef:.4f}')
Real intercept: -0.7370
Real coefficient for mean radius: 3.4229
Real coefficient for mean texture: 1.3558
Real coefficient for mean concavity: 1.2319
Real coefficient for mean smoothness: 1.3290

Each real coefficient here is a change in log-odds, not a change in probability directly; a coefficient of 3.42 for radius means each one real standard deviation increase in radius multiplies the real odds of malignancy by about e to the 3.42, roughly 30 times. That's a real, large effect, consistent with tumor size being one of the strongest real predictors of malignancy.

Part 4: The Confusion Matrix#

y_pred = model.predict(X_test_scaled)
cm = confusion_matrix(y_test, y_pred)
print('Real confusion matrix (rows=actual benign/malignant, cols=predicted benign/malignant):')
print(cm)
true_neg, false_pos, false_neg, true_pos = cm.ravel()
print(f'Real true negatives (correctly benign): {true_neg}')
print(f'Real false positives (benign called malignant): {false_pos}')
print(f'Real false negatives (malignant missed): {false_neg}')
print(f'Real true positives (correctly caught malignant): {true_pos}')
Real confusion matrix (rows=actual benign/malignant, cols=predicted benign/malignant):
[[86  4]
 [ 6 47]]
Real true negatives (correctly benign): 86
Real false positives (benign called malignant): 4
Real false negatives (malignant missed): 6
Real true positives (correctly caught malignant): 47

Part 5: Precision, Recall, and F1#

accuracy = accuracy_score(y_test, y_pred)
precision = precision_score(y_test, y_pred)
recall = recall_score(y_test, y_pred)
f1 = f1_score(y_test, y_pred)
print(f'Real accuracy: {accuracy:.4f}')
print(f'Real precision: {precision:.4f}')
print(f'Real recall: {recall:.4f}')
print(f'Real F1 score: {f1:.4f}')
Real accuracy: 0.9301
Real precision: 0.9216
Real recall: 0.8868
Real F1 score: 0.9038

In a real medical screening context, recall usually matters more than precision: a missed real cancer is far more costly than a false alarm that a follow-up test can rule out. This real model's default recall of 0.89 means about 11 percent of real malignant cases would be missed, which is exactly the kind of number worth trying to improve before precision.

Part 6: The Threshold Tradeoff#

y_proba = model.predict_proba(X_test_scaled)[:, 1]
for threshold in [0.5, 0.3, 0.2]:
    y_pred_t = (y_proba >= threshold).astype(int)
    p = precision_score(y_test, y_pred_t)
    r = recall_score(y_test, y_pred_t)
    cm_t = confusion_matrix(y_test, y_pred_t)
    print(f'threshold={threshold}: real precision={p:.4f}, real recall={r:.4f}, real false negatives={cm_t[1, 0]}, real false positives={cm_t[0, 1]}')
threshold=0.5: real precision=0.9216, real recall=0.8868, real false negatives=6, real false positives=4
threshold=0.3: real precision=0.8621, real recall=0.9434, real false negatives=3, real false positives=8
threshold=0.2: real precision=0.8226, real recall=0.9623, real false negatives=2, real false positives=11

Part 7: The ROC Curve and AUC#

fpr, tpr, thresholds = roc_curve(y_test, y_proba)
auc = roc_auc_score(y_test, y_proba)
print(f'Real AUC: {auc:.4f}')
plt.figure(figsize=(6, 6))
plt.plot(fpr, tpr, color='steelblue', label=f'Real model (AUC={auc:.3f})')
plt.plot([0, 1], [0, 1], color='gray', linestyle='--', label='Random guessing')
plt.title('Real ROC Curve')
plt.xlabel('False Positive Rate')
plt.ylabel('True Positive Rate')
plt.legend()
plt.show()
Real AUC: 0.9849
No description has been provided for this image

Part 8: Scaling Up to All 30 Real Features#

feature_cols = [c for c in cancer.columns if c not in ('target', 'malignant')]
X_full = cancer[feature_cols].values
X_train_full, X_test_full, y_train_full, y_test_full = train_test_split(X_full, y, test_size=0.25, random_state=42, stratify=y)
scaler_full = StandardScaler()
X_train_full_scaled = scaler_full.fit_transform(X_train_full)
X_test_full_scaled = scaler_full.transform(X_test_full)
full_model = LogisticRegression(max_iter=5000).fit(X_train_full_scaled, y_train_full)
full_auc = roc_auc_score(y_test_full, full_model.predict_proba(X_test_full_scaled)[:, 1])
print(f'Real 4-feature AUC: {auc:.4f}')
print(f'Real 30-feature AUC: {full_auc:.4f}')
print(f'Real 30-feature accuracy: {full_model.score(X_test_full_scaled, y_test_full):.4f}')
Real 4-feature AUC: 0.9849
Real 30-feature AUC: 0.9962
Real 30-feature accuracy: 0.9650

Part 9: Regularization Strength#

skf = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
for C in [0.01, 0.1, 1.0, 10.0, 100.0]:
    scores = cross_val_score(LogisticRegression(max_iter=5000, C=C), X_train_full_scaled, y_train_full, cv=skf, scoring='roc_auc')
    print(f'C={C}: real mean CV AUC={scores.mean():.4f}')
C=0.01: real mean CV AUC=0.9904
C=0.1: real mean CV AUC=0.9924
C=1.0: real mean CV AUC=0.9912
C=10.0: real mean CV AUC=0.9876
C=100.0: real mean CV AUC=0.9848

More real features didn't call for less regularization, it called for more. C equals 100 gave the model the most real freedom to fit the training data closely, and it generalized the worst of the five real settings tested. That's the same real theme from video one: the setting that fits training data best is not automatically the setting that predicts new real data best.

Part 10: Strengths, Weaknesses, and When to Use It#

  • Strength: outputs a genuine, calibrated real probability, not just a hard label.
  • Strength: coefficients are interpretable as real effects on log-odds.
  • Strength: the decision threshold can be tuned after training to match real-world costs, as part six showed.
  • Weakness: assumes a real linear decision boundary in log-odds space; genuinely non-linear real patterns need a different algorithm.
  • Use it as a real, fast, interpretable baseline for any binary classification problem before reaching for something more complex.

Part 11: Saving Your Work#

final_model = LogisticRegression(max_iter=5000, C=0.1).fit(X_train_full_scaled, y_train_full)
final_proba = final_model.predict_proba(X_test_full_scaled)[:, 1]
final_auc = roc_auc_score(y_test_full, final_proba)
final_fpr, final_tpr, _ = roc_curve(y_test_full, final_proba)
plt.figure(figsize=(6, 6))
plt.plot(final_fpr, final_tpr, color='steelblue', label=f'Final model (AUC={final_auc:.3f})')
plt.plot([0, 1], [0, 1], color='gray', linestyle='--')
plt.title('Real Final Model ROC Curve')
plt.xlabel('False Positive Rate')
plt.ylabel('True Positive Rate')
plt.legend()
plt.savefig('logistic_regression_final_roc.png', dpi=150)
plt.show()
No description has been provided for this image
import joblib
joblib.dump(final_model, 'logistic_regression_model.joblib')
reloaded_model = joblib.load('logistic_regression_model.joblib')
print(f'Real original model test accuracy: {final_model.score(X_test_full_scaled, y_test_full):.4f}')
print(f'Real reloaded model test accuracy: {reloaded_model.score(X_test_full_scaled, y_test_full):.4f}')
Real original model test accuracy: 0.9790
Real reloaded model test accuracy: 0.9790

Wrap-Up: What You Learned#

  • The real sigmoid function, and why it's the right tool for turning a linear score into a real probability.
  • Fitting real logistic regression and interpreting its coefficients as real log-odds effects.
  • The confusion matrix, and the real difference between precision, recall, and F1.
  • Tuning the real decision threshold to trade precision for recall, a genuinely practical, real-world lever.
  • The ROC curve and AUC as a threshold-free real summary of model quality.
  • Regularization strength, and a real case where the least-constrained model generalized worst.
  • Video three moves to K-Nearest Neighbors, a completely different kind of classifier, on the real Iris dataset. Subscribe so it lands automatically see you there.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.