Lesson 33 · Data Science Projects
Classifying Cancer Cells with Scikit-learn: A Step-by-Step Guide for Beginners
In this lesson, you will learn how to build a simple machine learning model to classify cancer cells. Accurate classification helps doctors detect if a cell…
- CourseData Science Projects
- Lesson33 of 33
- Video24 min
- FormatJupyter notebook · 15 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbCancer Cell Classification with Scikit learn#
- In this lesson, you will learn how to build a simple machine learning model to classify cancer cells.
- Accurate classification helps doctors detect if a cell is benign (not cancer) or malignant (cancer).
- Machine learning can help uncover useful signals in complex datasets with many measurements.
- You do not need any experience to follow along.
- We will use a popular dataset built into scikit learn, making it safe and easy.
- At the end, you will know how to train a model and check how well it performs.
What is Cancer Cell Classification?#
- Classification means teaching a computer program to assign a label, like benign or malignant, to examples based on their features.
- For cancer cell data, the model looks at measurements from cell samples and learns what patterns match each label.
- Good classification can save lives by helping with early detection and reducing human error.
# Always start by suppressing warnings for a cleaner notebook.
import warnings
import numpy as np
np.random.seed(42)
warnings.filterwarnings("ignore")
# Data setup: Breast Cancer Dataset from scikit learn
from sklearn.datasets import load_breast_cancer
data = load_breast_cancer()
X, y = data.data, data.target
print("Feature matrix shape:", X.shape)
print("Target label shape:", y.shape)
print("First five labels:", y[:5])
print("First five samples:")
print(X[:5])
# The target labels are 0 for malignant and 1 for benign.
print("Target names:", data.target_names)
print("Feature names:", data.feature_names)
print("Number of samples:", len(y))
Splitting Your Data for Training and Testing#
- To measure how well our machine learning model learns, we split our data into two sets.
- Training data is used by the model to learn, while testing data checks how well it performs on new examples.
- This split helps us avoid overfitting, which is when a model memorizes everything but cannot generalize.
# Let us split our dataset: 80% for training, 20% for testing.
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42
)
print("Training samples:", X_train.shape[0])
print("Testing samples:", X_test.shape[0])
# Data check: See if the classes are balanced in both sets.
import numpy as np
print("Training set label counts:", np.bincount(y_train))
print("Testing set label counts:", np.bincount(y_test))
# Optional: Standardize features so all measurements are on a similar scale.
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)
X_test_scaled = scaler.transform(X_test)
Choosing a Machine Learning Algorithm#
- There are many algorithms for classification. We will use Logistic Regression, which is simple and easy to interpret.
- Logistic Regression is commonly used for binary classification. It predicts the chance that a cell is malignant or benign.
- We will use scikit learn's built in implementation.
# Train a Logistic Regression classifier.
from sklearn.linear_model import LogisticRegression
clf = LogisticRegression(max_iter=10000, random_state=42)
clf.fit(X_train_scaled, y_train)
# Check the accuracy of the model on training and test data.
train_accuracy = clf.score(X_train_scaled, y_train)
test_accuracy = clf.score(X_test_scaled, y_test)
print(f"Training accuracy: {train_accuracy:.2f}")
print(f"Test accuracy: {test_accuracy:.2f}")
# Look deeper: Show a confusion matrix and a classification report.
from sklearn.metrics import confusion_matrix, classification_report
y_pred = clf.predict(X_test_scaled)
cm = confusion_matrix(y_test, y_pred)
cr = classification_report(y_test, y_pred, target_names=data.target_names)
print("Confusion Matrix:\n", cm)
print("\nClassification Report:\n", cr)
# Visualize: Plot a few features to understand separation.
import matplotlib.pyplot as plt
plt.figure(figsize=(7,5))
plt.scatter(X_train_scaled[:, 0], X_train_scaled[:, 1], c=y_train, cmap='coolwarm', alpha=0.6)
plt.xlabel(data.feature_names[0])
plt.ylabel(data.feature_names[1])
plt.title("Cancer cell samples (first two features)")
plt.show()
Trying Other Models: Support Vector Machine#
- You can try other classifiers, like Support Vector Machine (SVM).
- SVM tries to find the best boundary that separates the two classes.
- Changing algorithms helps you see which method works best for your data.
# Train and test a Support Vector Machine (SVM) classifier.
from sklearn.svm import SVC
svm_model = SVC(kernel='linear', random_state=42)
svm_model.fit(X_train_scaled, y_train)
svm_test_acc = svm_model.score(X_test_scaled, y_test)
print(f"SVM test accuracy: {svm_test_acc:.2f}")
# Bonus: Feature importance with logistic regression coefficients.
import pandas as pd
coef = clf.coef_[0]
importance = pd.Series(coef, index=data.feature_names)
print(importance.sort_values(key=abs, ascending=False).head(10))
# Make a prediction for a single sample (from the test set).
sample = X_test_scaled[0]
true_label = y_test[0]
predicted_label = clf.predict([sample])[0]
print(f"True label: {data.target_names[true_label]}")
print(f"Predicted label: {data.target_names[predicted_label]}")
# You try: Predict for a new (simulated) sample.
import numpy as np
user_sample = np.mean(X_train, axis=0) + np.random.randn(X_train.shape[1]) * 0.1
user_sample_scaled = scaler.transform([user_sample])
prediction = clf.predict(user_sample_scaled)[0]
print(f"Predicted label for your made up sample: {data.target_names[prediction]}")
# Mini project: Ask the user for measurements and predict cancer cell type.
user_features = []
for name in data.feature_names[:5]:
val = float(input(f"Enter value for {name}: "))
user_features.append(val)
user_features = user_features + list(np.mean(X_train, axis=0)[5:])
user_features_scaled = scaler.transform([user_features])
prediction = clf.predict(user_features_scaled)[0]
print(f"Prediction: {data.target_names[prediction]}")
What You Learned#
- You loaded and explored a real life cancer dataset.
- You prepared your data for learning and visualized it.
- You built and evaluated a classifier using scikit learn.
- You checked model performance using several metrics.
- You even tried a small project with your own values.
- Try substituting other scikit learn datasets and algorithms to keep learning.
- Subscribe to our YouTube channel for more beginner friendly data science lessons!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



