Mathew K Analytics

Lesson 33 · Data Science Projects

Classifying Cancer Cells with Scikit-learn: A Step-by-Step Guide for Beginners

In this lesson, you will learn how to build a simple machine learning model to classify cancer cells. Accurate classification helps doctors detect if a cell…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Cancer Cell Classification with Scikit learn#

  • In this lesson, you will learn how to build a simple machine learning model to classify cancer cells.
  • Accurate classification helps doctors detect if a cell is benign (not cancer) or malignant (cancer).
  • Machine learning can help uncover useful signals in complex datasets with many measurements.
  • You do not need any experience to follow along.
  • We will use a popular dataset built into scikit learn, making it safe and easy.
  • At the end, you will know how to train a model and check how well it performs.

What is Cancer Cell Classification?#

  • Classification means teaching a computer program to assign a label, like benign or malignant, to examples based on their features.
  • For cancer cell data, the model looks at measurements from cell samples and learns what patterns match each label.
  • Good classification can save lives by helping with early detection and reducing human error.
# Always start by suppressing warnings for a cleaner notebook.
import warnings
import numpy as np
np.random.seed(42)
warnings.filterwarnings("ignore")
 
# Data setup: Breast Cancer Dataset from scikit learn
from sklearn.datasets import load_breast_cancer
data = load_breast_cancer()
X, y = data.data, data.target
print("Feature matrix shape:", X.shape)
print("Target label shape:", y.shape)
print("First five labels:", y[:5])
print("First five samples:")
print(X[:5])
Feature matrix shape: (569, 30)
Target label shape: (569,)
First five labels: [0 0 0 0 0]
First five samples:
[[1.799e+01 1.038e+01 1.228e+02 1.001e+03 1.184e-01 2.776e-01 3.001e-01
  1.471e-01 2.419e-01 7.871e-02 1.095e+00 9.053e-01 8.589e+00 1.534e+02
  6.399e-03 4.904e-02 5.373e-02 1.587e-02 3.003e-02 6.193e-03 2.538e+01
  1.733e+01 1.846e+02 2.019e+03 1.622e-01 6.656e-01 7.119e-01 2.654e-01
  4.601e-01 1.189e-01]
 [2.057e+01 1.777e+01 1.329e+02 1.326e+03 8.474e-02 7.864e-02 8.690e-02
  7.017e-02 1.812e-01 5.667e-02 5.435e-01 7.339e-01 3.398e+00 7.408e+01
  5.225e-03 1.308e-02 1.860e-02 1.340e-02 1.389e-02 3.532e-03 2.499e+01
  2.341e+01 1.588e+02 1.956e+03 1.238e-01 1.866e-01 2.416e-01 1.860e-01
  2.750e-01 8.902e-02]
 [1.969e+01 2.125e+01 1.300e+02 1.203e+03 1.096e-01 1.599e-01 1.974e-01
  1.279e-01 2.069e-01 5.999e-02 7.456e-01 7.869e-01 4.585e+00 9.403e+01
  6.150e-03 4.006e-02 3.832e-02 2.058e-02 2.250e-02 4.571e-03 2.357e+01
  2.553e+01 1.525e+02 1.709e+03 1.444e-01 4.245e-01 4.504e-01 2.430e-01
  3.613e-01 8.758e-02]
 [1.142e+01 2.038e+01 7.758e+01 3.861e+02 1.425e-01 2.839e-01 2.414e-01
  1.052e-01 2.597e-01 9.744e-02 4.956e-01 1.156e+00 3.445e+00 2.723e+01
  9.110e-03 7.458e-02 5.661e-02 1.867e-02 5.963e-02 9.208e-03 1.491e+01
  2.650e+01 9.887e+01 5.677e+02 2.098e-01 8.663e-01 6.869e-01 2.575e-01
  6.638e-01 1.730e-01]
 [2.029e+01 1.434e+01 1.351e+02 1.297e+03 1.003e-01 1.328e-01 1.980e-01
  1.043e-01 1.809e-01 5.883e-02 7.572e-01 7.813e-01 5.438e+00 9.444e+01
  1.149e-02 2.461e-02 5.688e-02 1.885e-02 1.756e-02 5.115e-03 2.254e+01
  1.667e+01 1.522e+02 1.575e+03 1.374e-01 2.050e-01 4.000e-01 1.625e-01
  2.364e-01 7.678e-02]]
# The target labels are 0 for malignant and 1 for benign.
print("Target names:", data.target_names)
print("Feature names:", data.feature_names)
print("Number of samples:", len(y))
Target names: ['malignant' 'benign']
Feature names: ['mean radius' 'mean texture' 'mean perimeter' 'mean area'
 'mean smoothness' 'mean compactness' 'mean concavity'
 'mean concave points' 'mean symmetry' 'mean fractal dimension'
 'radius error' 'texture error' 'perimeter error' 'area error'
 'smoothness error' 'compactness error' 'concavity error'
 'concave points error' 'symmetry error' 'fractal dimension error'
 'worst radius' 'worst texture' 'worst perimeter' 'worst area'
 'worst smoothness' 'worst compactness' 'worst concavity'
 'worst concave points' 'worst symmetry' 'worst fractal dimension']
Number of samples: 569

Splitting Your Data for Training and Testing#

  • To measure how well our machine learning model learns, we split our data into two sets.
  • Training data is used by the model to learn, while testing data checks how well it performs on new examples.
  • This split helps us avoid overfitting, which is when a model memorizes everything but cannot generalize.
# Let us split our dataset: 80% for training, 20% for testing.
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42
)
print("Training samples:", X_train.shape[0])
print("Testing samples:", X_test.shape[0])
Training samples: 455
Testing samples: 114
# Data check: See if the classes are balanced in both sets.
import numpy as np
print("Training set label counts:", np.bincount(y_train))
print("Testing set label counts:", np.bincount(y_test))
Training set label counts: [169 286]
Testing set label counts: [43 71]
# Optional: Standardize features so all measurements are on a similar scale.
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)
X_test_scaled = scaler.transform(X_test)

Choosing a Machine Learning Algorithm#

  • There are many algorithms for classification. We will use Logistic Regression, which is simple and easy to interpret.
  • Logistic Regression is commonly used for binary classification. It predicts the chance that a cell is malignant or benign.
  • We will use scikit learn's built in implementation.
# Train a Logistic Regression classifier.
from sklearn.linear_model import LogisticRegression
clf = LogisticRegression(max_iter=10000, random_state=42)
clf.fit(X_train_scaled, y_train)
 
LogisticRegression(max_iter=10000, random_state=42)
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
# Check the accuracy of the model on training and test data.
train_accuracy = clf.score(X_train_scaled, y_train)
test_accuracy = clf.score(X_test_scaled, y_test)
print(f"Training accuracy: {train_accuracy:.2f}")
print(f"Test accuracy: {test_accuracy:.2f}")
Training accuracy: 0.99
Test accuracy: 0.97
# Look deeper: Show a confusion matrix and a classification report.
from sklearn.metrics import confusion_matrix, classification_report
y_pred = clf.predict(X_test_scaled)
cm = confusion_matrix(y_test, y_pred)
cr = classification_report(y_test, y_pred, target_names=data.target_names)
print("Confusion Matrix:\n", cm)
print("\nClassification Report:\n", cr)
Confusion Matrix:
 [[41  2]
 [ 1 70]]

Classification Report:
               precision    recall  f1-score   support

   malignant       0.98      0.95      0.96        43
      benign       0.97      0.99      0.98        71

    accuracy                           0.97       114
   macro avg       0.97      0.97      0.97       114
weighted avg       0.97      0.97      0.97       114

# Visualize: Plot a few features to understand separation.
import matplotlib.pyplot as plt
plt.figure(figsize=(7,5))
plt.scatter(X_train_scaled[:, 0], X_train_scaled[:, 1], c=y_train, cmap='coolwarm', alpha=0.6)
plt.xlabel(data.feature_names[0])
plt.ylabel(data.feature_names[1])
plt.title("Cancer cell samples (first two features)")
plt.show()
No description has been provided for this image

Trying Other Models: Support Vector Machine#

  • You can try other classifiers, like Support Vector Machine (SVM).
  • SVM tries to find the best boundary that separates the two classes.
  • Changing algorithms helps you see which method works best for your data.
# Train and test a Support Vector Machine (SVM) classifier.
from sklearn.svm import SVC
svm_model = SVC(kernel='linear', random_state=42)
svm_model.fit(X_train_scaled, y_train)
svm_test_acc = svm_model.score(X_test_scaled, y_test)
print(f"SVM test accuracy: {svm_test_acc:.2f}")
SVM test accuracy: 0.96
# Bonus: Feature importance with logistic regression coefficients.
import pandas as pd
coef = clf.coef_[0]
importance = pd.Series(coef, index=data.feature_names)
print(importance.sort_values(key=abs, ascending=False).head(10))
worst texture          -1.350606
radius error           -1.268178
worst symmetry         -1.208200
mean concave points    -1.119804
worst concavity        -0.943053
area error             -0.907186
worst radius           -0.879840
worst area             -0.841846
mean concavity         -0.801458
worst concave points   -0.778217
dtype: float64
# Make a prediction for a single sample (from the test set).
sample = X_test_scaled[0]
true_label = y_test[0]
predicted_label = clf.predict([sample])[0]
print(f"True label: {data.target_names[true_label]}")
print(f"Predicted label: {data.target_names[predicted_label]}")
True label: benign
Predicted label: benign
# You try: Predict for a new (simulated) sample.
import numpy as np
user_sample = np.mean(X_train, axis=0) + np.random.randn(X_train.shape[1]) * 0.1
user_sample_scaled = scaler.transform([user_sample])
prediction = clf.predict(user_sample_scaled)[0]
print(f"Predicted label for your made up sample: {data.target_names[prediction]}")
Predicted label for your made up sample: malignant
# Mini project: Ask the user for measurements and predict cancer cell type.
user_features = []
for name in data.feature_names[:5]:
    val = float(input(f"Enter value for {name}: "))
    user_features.append(val)
user_features = user_features + list(np.mean(X_train, axis=0)[5:])
user_features_scaled = scaler.transform([user_features])
prediction = clf.predict(user_features_scaled)[0]
print(f"Prediction: {data.target_names[prediction]}")
Prediction: malignant

What You Learned#

  • You loaded and explored a real life cancer dataset.
  • You prepared your data for learning and visualized it.
  • You built and evaluated a classifier using scikit learn.
  • You checked model performance using several metrics.
  • You even tried a small project with your own values.
  • Try substituting other scikit learn datasets and algorithms to keep learning.
  • Subscribe to our YouTube channel for more beginner friendly data science lessons!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.