Mathew K Analytics

Lesson 16 · Data Mining

Understanding Decision Trees: ID3, C4.5, and CART Algorithms for Classification

Welcome! In this hands-on lesson, we will explore how decision trees work for real-world data mining tasks. We will use the Telecom Customer Churn and…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Decision Trees in Python (ID3, C4.5, CART): Telecom Churn & Titanic#

Welcome! In this hands-on lesson, we will explore how decision trees work for real-world data mining tasks.

We will use the Telecom Customer Churn and Titanic datasets. By the end, you will be able to build, understand, and evaluate your own decision tree models.

Let's get started!

import warnings
import numpy as np
np.random.seed(42)
warnings.filterwarnings("ignore")

# Suppressing warning messages helps beginners focus.

What is a Decision Tree?#

A decision tree is a way to make predictions by splitting data into smaller and smaller groups. Picture a set of yes-or-no questions that lead to a final answer.

In data mining, decision trees are used for classification (predicting labels) and regression (predicting numbers).

# Data setup (Telecom Customer Churn Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/IBM/telco-customer-churn-on-icp4d/master/data/Telco-Customer-Churn.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
(7043, 21)
   customerID  gender  SeniorCitizen Partner Dependents  tenure PhoneService  \
0  7590-VHVEG  Female              0     Yes         No       1           No   
1  5575-GNVDE    Male              0      No         No      34          Yes   
2  3668-QPYBK    Male              0      No         No       2          Yes   

      MultipleLines InternetService OnlineSecurity  ... DeviceProtection  \
0  No phone service             DSL             No  ...               No   
1                No             DSL            Yes  ...              Yes   
2                No             DSL            Yes  ...               No   

  TechSupport StreamingTV StreamingMovies        Contract PaperlessBilling  \
0          No          No              No  Month-to-month              Yes   
1          No          No              No        One year               No   
2          No          No              No  Month-to-month              Yes   

      PaymentMethod MonthlyCharges  TotalCharges Churn  
0  Electronic check          29.85         29.85    No  
1      Mailed check          56.95        1889.5    No  
2      Mailed check          53.85        108.15   Yes  

[3 rows x 21 columns]
# Basic data cleaning
df = df.copy()
df = df.drop(['customerID'], axis=1)
df['TotalCharges'] = pd.to_numeric(df['TotalCharges'], errors='coerce')
missing = df.isnull().sum().sum()
print('Missing values:', missing)
df = df.dropna()
print(df.shape)
Missing values: 11
(7032, 20)

Turning categorical variables into numbers#

Decision trees in Python like numbers, not text. We will convert Yes/No columns and other text into 1s and 0s.

# Convert object columns to categorical numbers
for col in df.columns:
    if df[col].dtype == 'object':
        if set(df[col].unique()) == set(['Yes', 'No']):
            df[col] = df[col].map({'Yes': 1, 'No': 0})
        else:
            df[col] = pd.factorize(df[col])[0]
print(df.dtypes.head())
gender           int64
SeniorCitizen    int64
Partner          int64
Dependents       int64
tenure           int64
dtype: object

Exploratory Data Analysis (EDA) Preview#

Let us take a quick look at a few features most related to churn, the business outcome.

This helps us see trends before modeling.

import matplotlib.pyplot as plt
import seaborn as sns
sns.set(style='whitegrid')

# Boxplot: MonthlyCharges by Churn
plt.figure(figsize=(6,4))
sns.boxplot(x='Churn', y='MonthlyCharges', data=df)
plt.title('Boxplot of Monthly Charges by Churn Status')
plt.show()

# Countplot: Number of Churned customers
sns.countplot(x='Churn', data=df)
plt.title('Churn Counts')
plt.show()
No description has been provided for this image
No description has been provided for this image

Decision Tree: Theory Pulse#

Decision trees use splitting rules based on the data. Popular ones are:

  • ID3: Uses information gain (entropy)
  • C4.5: Improves ID3, uses gain ratio
  • CART: Uses Gini index, can do regression too

These all create flowchart-like models that are easy to explain.

# Train-test split
from sklearn.model_selection import train_test_split
target = 'Churn'
X = df.drop(target, axis=1)
y = df[target]

X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, random_state=42)
print(X_train.shape, X_test.shape)
(5274, 19) (1758, 19)
# CART Decision Tree (scikit-learn)
from sklearn.tree import DecisionTreeClassifier
tree_cart = DecisionTreeClassifier(criterion='gini', random_state=42, max_depth=4)
tree_cart.fit(X_train, y_train)
print('Train accuracy:', tree_cart.score(X_train, y_train))
print('Test accuracy:', tree_cart.score(X_test, y_test))
Train accuracy: 0.7969283276450512
Test accuracy: 0.7804323094425484
# Visualize the CART tree
from sklearn.tree import plot_tree
plt.figure(figsize=(18,10))
plot_tree(tree_cart, feature_names=X_train.columns, class_names=['No','Yes'], filled=True, max_depth=3)
plt.show()
No description has been provided for this image
# Confusion Matrix and Classification Report
from sklearn.metrics import confusion_matrix, classification_report
y_pred = tree_cart.predict(X_test)
print('Confusion Matrix:')
print(confusion_matrix(y_test, y_pred))
print('Classification Report:')
print(classification_report(y_test, y_pred))
Confusion Matrix:
[[1163  137]
 [ 249  209]]
Classification Report:
              precision    recall  f1-score   support

           0       0.82      0.89      0.86      1300
           1       0.60      0.46      0.52       458

    accuracy                           0.78      1758
   macro avg       0.71      0.68      0.69      1758
weighted avg       0.77      0.78      0.77      1758

# Try Decision Tree ID3 (entropy)
tree_id3 = DecisionTreeClassifier(criterion='entropy', random_state=42, max_depth=4)
tree_id3.fit(X_train, y_train)
print('Train accuracy (ID3):', tree_id3.score(X_train, y_train))
print('Test accuracy (ID3):', tree_id3.score(X_test, y_test))
Train accuracy (ID3): 0.7967387182404247
Test accuracy (ID3): 0.7815699658703071
# C4.5-style: using criterion='entropy' + pruning
tree_c45 = DecisionTreeClassifier(criterion='entropy', max_depth=4, min_samples_split=20, random_state=42)
tree_c45.fit(X_train, y_train)
print('Train accuracy (C4.5 style):', tree_c45.score(X_train, y_train))
print('Test accuracy (C4.5 style):', tree_c45.score(X_test, y_test))
Train accuracy (C4.5 style): 0.7967387182404247
Test accuracy (C4.5 style): 0.7815699658703071
# Feature importance for the tree
import numpy as np
importances = tree_cart.feature_importances_
indices = np.argsort(importances)[::-1][:5]
features = X_train.columns[indices]
print('Top 5 features for CART:')
for i in range(5):
    print(f"{features[i]}: {importances[indices[i]]:.3f}")
    
Top 5 features for CART:
Contract: 0.570
MonthlyCharges: 0.163
tenure: 0.162
TotalCharges: 0.041
OnlineSecurity: 0.029

Mini-Project: Try the Titanic Dataset#

Now, let's predict survival on the Titanic using what we've learned.

This will help you see if your skills work across more than one problem.

# Data setup (Titanic Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
df2 = pd.read_csv(url)
print(df2.shape)
print(df2.head(3))
(891, 12)
   PassengerId  Survived  Pclass  \
0            1         0       3   
1            2         1       1   
2            3         1       3   

                                                Name     Sex   Age  SibSp  \
0                            Braund, Mr. Owen Harris    male  22.0      1   
1  Cumings, Mrs. John Bradley (Florence Briggs Th...  female  38.0      1   
2                             Heikkinen, Miss. Laina  female  26.0      0   

   Parch            Ticket     Fare Cabin Embarked  
0      0         A/5 21171   7.2500   NaN        S  
1      0          PC 17599  71.2833   C85        C  
2      0  STON/O2. 3101282   7.9250   NaN        S  
# Prepare Titanic data for decision trees
df2 = df2[['Survived', 'Pclass', 'Sex', 'Age', 'SibSp', 'Parch', 'Fare', 'Embarked']].copy()
df2['Age'] = df2['Age'].fillna(df2['Age'].median())
df2['Embarked'] = df2['Embarked'].fillna('S')
df2['Sex'] = df2['Sex'].map({'male': 0, 'female': 1})
df2['Embarked'] = pd.factorize(df2['Embarked'])[0]
# Fit a Titanic tree
from sklearn.model_selection import train_test_split
X2 = df2.drop('Survived', axis=1)
y2 = df2['Survived']
X2_train, X2_test, y2_train, y2_test = train_test_split(X2, y2, test_size=0.25, random_state=42)
tree_titanic = DecisionTreeClassifier(max_depth=3, random_state=42)
tree_titanic.fit(X2_train, y2_train)
print('Train accuracy:', tree_titanic.score(X2_train, y2_train))
print('Test accuracy:', tree_titanic.score(X2_test, y2_test))
Train accuracy: 0.8323353293413174
Test accuracy: 0.8026905829596412
# Interactive: Your prediction
pclass = int(input("Enter passenger class (1, 2, or 3): "))
sex = int(input("Enter sex (0 = male, 1 = female): "))
age = float(input("Enter age: "))
sibsp = int(input("Enter number of siblings/spouses: "))
parch = int(input("Enter number of parents/children: "))
fare = float(input("Enter fare: "))
embarked = int(input("Enter port (0 = S, 1 = C, 2 = Q): "))
sample = [[pclass, sex, age, sibsp, parch, fare, embarked]]
prediction = tree_titanic.predict(sample)[0]
print("Prediction (1 = survived, 0 = did not survive):", prediction)
Prediction (1 = survived, 0 = did not survive): 0

Best Practices and Troubleshooting Tips#

  1. Always check if there are missing values after you load data.
  2. Visualize features to look for outliers or surprises.
  3. Set random_state in train_test_split for stable results.
  4. Try different max_depth, min_samples_split, and split criteria.
  5. Avoid leaking the answer in your features.

If your tree overfits, prune it or require more samples at each branch.

Extra Decision Tree Challenges!#

  • Challenge 1: Tune your tree's depth and report a test score.
  • Challenge 2: Use feature_importances_ to pick the top 3 variables and build a mini-tree.
  • Challenge 3: Try predicting a friend's survival on Titanic using your tree.

Bonus: Show your plot as a tree diagram, and try explaining a branch in plain language.

Recap: What Did We Learn?#

  • What decision trees are and why they're great for beginners.
  • The difference between ID3, C4.5, and CART.
  • How to clean, prepare, and split data in pandas.
  • How to fit and evaluate trees in scikit-learn.
  • How to read feature importance and explain results.

Congratulations, you have now built and interpreted your own decision trees in Python!

Thanks for Joining Next Steps!#

Keep practicing building decision trees. Try out different datasets or challenges from our playlist.

If you liked this lesson, give it a like and subscribe to our YouTube channel for more hands-on data mining guides!

See you next time!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.