Mathew K Analytics

Lesson 19 · Data Mining

Master Logistic Regression for Binary Classification: Step-by-Step Practical Guide

Welcome! This week we explore Logistic Regression for classification problems. Logistic Regression predicts categories, like: Will a phone customer churn?…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Week 56: Logistic Regression for Classification#

Welcome! This week we explore Logistic Regression for classification problems. Logistic Regression predicts categories, like: Will a phone customer churn?

We use the Telecom Customer Churn Dataset. Let's dive in!

# Suppress warnings for clean output
import warnings; warnings.filterwarnings("ignore")
import numpy as np
np.random.seed(42)
# Data setup (Telecom Customer Churn Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/IBM/telco-customer-churn-on-icp4d/master/data/Telco-Customer-Churn.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
(7043, 21)
   customerID  gender  SeniorCitizen Partner Dependents  tenure PhoneService  \
0  7590-VHVEG  Female              0     Yes         No       1           No   
1  5575-GNVDE    Male              0      No         No      34          Yes   
2  3668-QPYBK    Male              0      No         No       2          Yes   

      MultipleLines InternetService OnlineSecurity  ... DeviceProtection  \
0  No phone service             DSL             No  ...               No   
1                No             DSL            Yes  ...              Yes   
2                No             DSL            Yes  ...               No   

  TechSupport StreamingTV StreamingMovies        Contract PaperlessBilling  \
0          No          No              No  Month-to-month              Yes   
1          No          No              No        One year               No   
2          No          No              No  Month-to-month              Yes   

      PaymentMethod MonthlyCharges  TotalCharges Churn  
0  Electronic check          29.85         29.85    No  
1      Mailed check          56.95        1889.5    No  
2      Mailed check          53.85        108.15   Yes  

[3 rows x 21 columns]

What is Logistic Regression?#

Logistic Regression helps us predict a category. It is great for problems like: Is a customer likely to leave (churn) or stay?

The model gives us probabilities, and predicts yes or no.

# Take a quick look at how many customers churned
df['Churn'].value_counts()
Churn
No     5174
Yes    1869
Name: count, dtype: int64
# Check for missing data
df.isnull().sum().sort_values(ascending=False).head(5)
customerID       0
gender           0
SeniorCitizen    0
Partner          0
Dependents       0
dtype: int64
# See column data types
df.dtypes.head(8)
customerID       object
gender           object
SeniorCitizen     int64
Partner          object
Dependents       object
tenure            int64
PhoneService     object
MultipleLines    object
dtype: object
# Clean up TotalCharges  convert to numeric, fill missing values with median
df['TotalCharges'] = pd.to_numeric(df['TotalCharges'], errors='coerce')
df['TotalCharges'].fillna(df['TotalCharges'].median(), inplace=True)
# Quick stats for numeric columns
df.describe()
SeniorCitizen tenure MonthlyCharges TotalCharges
count 7043.000000 7043.000000 7043.000000 7043.000000
mean 0.162147 32.371149 64.761692 2281.916928
std 0.368612 24.559481 30.090047 2265.270398
min 0.000000 0.000000 18.250000 18.800000
25% 0.000000 9.000000 35.500000 402.225000
50% 0.000000 29.000000 70.350000 1397.475000
75% 0.000000 55.000000 89.850000 3786.600000
max 1.000000 72.000000 118.750000 8684.800000
# Convert categorical columns to numbers
df['Churn'] = df['Churn'].map({'Yes': 1, 'No': 0})
df['gender'] = df['gender'].map({'Male': 1, 'Female': 0})
# Pick features for our model
features = ['tenure', 'MonthlyCharges', 'TotalCharges', 'gender']
X = df[features]
y = df['Churn']
# Split data into training and test sets
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
# Create and train Logistic Regression model
from sklearn.linear_model import LogisticRegression
model = LogisticRegression(max_iter=200, solver='liblinear')
model.fit(X_train, y_train)
LogisticRegression(max_iter=200, solver='liblinear')
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
# Predict churn on the test set
y_pred = model.predict(X_test)
print(y_pred[:10])
[0 0 0 1 0 0 0 0 0 0]
# Evaluate accuracy
from sklearn.metrics import accuracy_score
accuracy = accuracy_score(y_test, y_pred)
print("Accuracy:", accuracy)
Accuracy: 0.7906316536550745
# See model probabilities
y_proba = model.predict_proba(X_test)
print(y_proba[:5])
[[0.57525199 0.42474801]
 [0.97238488 0.02761512]
 [0.98630201 0.01369799]
 [0.3216709  0.6783291 ]
 [0.9878224  0.0121776 ]]
# Visualize results: plot confusion matrix
import matplotlib.pyplot as plt
from sklearn.metrics import ConfusionMatrixDisplay
fig, ax = plt.subplots(figsize=(5, 5))
ConfusionMatrixDisplay.from_estimator(model, X_test, y_test, ax=ax)
plt.title('Churn Prediction: Confusion Matrix')
plt.show()
No description has been provided for this image

Try it yourself! Practice predicting with new values#

Can you make a prediction for a new customer with the model? What values would you use?

# Input your own example and predict
print("Enter tenure (months): ")
tenure = int(input())
print("Enter monthly charges: ")
monthly = float(input())
print("Enter total charges: ")
total = float(input())
print("Enter gender (1 for male, 0 for female): ")
gender = int(input())
sample = [[tenure, monthly, total, gender]]
prediction = model.predict(sample)
print("Predicted churn (1=Yes, 0=No): ", prediction[0])
Enter tenure (months): 
Enter monthly charges: 
Enter total charges: 
Enter gender (1 for male, 0 for female): 
Predicted churn (1=Yes, 0=No):  0

Recap & Whats Next#

This week, you learned to clean data, choose features, train and test a logistic regression model, and interpret results. Practice makes perfect! Experiment with other datasets and try different features.

Next: More on classification models, and fine-tuning your results.

Challenge: Change the Features#

Try changing which columns are used as features. See if your predictions and model accuracy change.

Optional: Use a different dataset and compare results.

Thank you for learning with us! To keep going, subscribe and join our next lesson!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.