Lesson 16 · Data Mining
Understanding Decision Trees: ID3, C4.5, and CART Algorithms for Classification
Welcome! In this hands-on lesson, we will explore how decision trees work for real-world data mining tasks. We will use the Telecom Customer Churn and…
- CourseData Mining
- Lesson16 of 31
- Video26 min
- FormatJupyter notebook · 16 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbDecision Trees in Python (ID3, C4.5, CART): Telecom Churn & Titanic#
Welcome! In this hands-on lesson, we will explore how decision trees work for real-world data mining tasks.
We will use the Telecom Customer Churn and Titanic datasets. By the end, you will be able to build, understand, and evaluate your own decision tree models.
Let's get started!
import warnings
import numpy as np
np.random.seed(42)
warnings.filterwarnings("ignore")
# Suppressing warning messages helps beginners focus.
What is a Decision Tree?#
A decision tree is a way to make predictions by splitting data into smaller and smaller groups. Picture a set of yes-or-no questions that lead to a final answer.
In data mining, decision trees are used for classification (predicting labels) and regression (predicting numbers).
# Data setup (Telecom Customer Churn Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/IBM/telco-customer-churn-on-icp4d/master/data/Telco-Customer-Churn.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
# Basic data cleaning
df = df.copy()
df = df.drop(['customerID'], axis=1)
df['TotalCharges'] = pd.to_numeric(df['TotalCharges'], errors='coerce')
missing = df.isnull().sum().sum()
print('Missing values:', missing)
df = df.dropna()
print(df.shape)
Turning categorical variables into numbers#
Decision trees in Python like numbers, not text. We will convert Yes/No columns and other text into 1s and 0s.
# Convert object columns to categorical numbers
for col in df.columns:
if df[col].dtype == 'object':
if set(df[col].unique()) == set(['Yes', 'No']):
df[col] = df[col].map({'Yes': 1, 'No': 0})
else:
df[col] = pd.factorize(df[col])[0]
print(df.dtypes.head())
Exploratory Data Analysis (EDA) Preview#
Let us take a quick look at a few features most related to churn, the business outcome.
This helps us see trends before modeling.
import matplotlib.pyplot as plt
import seaborn as sns
sns.set(style='whitegrid')
# Boxplot: MonthlyCharges by Churn
plt.figure(figsize=(6,4))
sns.boxplot(x='Churn', y='MonthlyCharges', data=df)
plt.title('Boxplot of Monthly Charges by Churn Status')
plt.show()
# Countplot: Number of Churned customers
sns.countplot(x='Churn', data=df)
plt.title('Churn Counts')
plt.show()
Decision Tree: Theory Pulse#
Decision trees use splitting rules based on the data. Popular ones are:
- ID3: Uses information gain (entropy)
- C4.5: Improves ID3, uses gain ratio
- CART: Uses Gini index, can do regression too
These all create flowchart-like models that are easy to explain.
# Train-test split
from sklearn.model_selection import train_test_split
target = 'Churn'
X = df.drop(target, axis=1)
y = df[target]
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, random_state=42)
print(X_train.shape, X_test.shape)
# CART Decision Tree (scikit-learn)
from sklearn.tree import DecisionTreeClassifier
tree_cart = DecisionTreeClassifier(criterion='gini', random_state=42, max_depth=4)
tree_cart.fit(X_train, y_train)
print('Train accuracy:', tree_cart.score(X_train, y_train))
print('Test accuracy:', tree_cart.score(X_test, y_test))
# Visualize the CART tree
from sklearn.tree import plot_tree
plt.figure(figsize=(18,10))
plot_tree(tree_cart, feature_names=X_train.columns, class_names=['No','Yes'], filled=True, max_depth=3)
plt.show()
# Confusion Matrix and Classification Report
from sklearn.metrics import confusion_matrix, classification_report
y_pred = tree_cart.predict(X_test)
print('Confusion Matrix:')
print(confusion_matrix(y_test, y_pred))
print('Classification Report:')
print(classification_report(y_test, y_pred))
# Try Decision Tree ID3 (entropy)
tree_id3 = DecisionTreeClassifier(criterion='entropy', random_state=42, max_depth=4)
tree_id3.fit(X_train, y_train)
print('Train accuracy (ID3):', tree_id3.score(X_train, y_train))
print('Test accuracy (ID3):', tree_id3.score(X_test, y_test))
# C4.5-style: using criterion='entropy' + pruning
tree_c45 = DecisionTreeClassifier(criterion='entropy', max_depth=4, min_samples_split=20, random_state=42)
tree_c45.fit(X_train, y_train)
print('Train accuracy (C4.5 style):', tree_c45.score(X_train, y_train))
print('Test accuracy (C4.5 style):', tree_c45.score(X_test, y_test))
# Feature importance for the tree
import numpy as np
importances = tree_cart.feature_importances_
indices = np.argsort(importances)[::-1][:5]
features = X_train.columns[indices]
print('Top 5 features for CART:')
for i in range(5):
print(f"{features[i]}: {importances[indices[i]]:.3f}")
Mini-Project: Try the Titanic Dataset#
Now, let's predict survival on the Titanic using what we've learned.
This will help you see if your skills work across more than one problem.
# Data setup (Titanic Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
df2 = pd.read_csv(url)
print(df2.shape)
print(df2.head(3))
# Prepare Titanic data for decision trees
df2 = df2[['Survived', 'Pclass', 'Sex', 'Age', 'SibSp', 'Parch', 'Fare', 'Embarked']].copy()
df2['Age'] = df2['Age'].fillna(df2['Age'].median())
df2['Embarked'] = df2['Embarked'].fillna('S')
df2['Sex'] = df2['Sex'].map({'male': 0, 'female': 1})
df2['Embarked'] = pd.factorize(df2['Embarked'])[0]
# Fit a Titanic tree
from sklearn.model_selection import train_test_split
X2 = df2.drop('Survived', axis=1)
y2 = df2['Survived']
X2_train, X2_test, y2_train, y2_test = train_test_split(X2, y2, test_size=0.25, random_state=42)
tree_titanic = DecisionTreeClassifier(max_depth=3, random_state=42)
tree_titanic.fit(X2_train, y2_train)
print('Train accuracy:', tree_titanic.score(X2_train, y2_train))
print('Test accuracy:', tree_titanic.score(X2_test, y2_test))
# Interactive: Your prediction
pclass = int(input("Enter passenger class (1, 2, or 3): "))
sex = int(input("Enter sex (0 = male, 1 = female): "))
age = float(input("Enter age: "))
sibsp = int(input("Enter number of siblings/spouses: "))
parch = int(input("Enter number of parents/children: "))
fare = float(input("Enter fare: "))
embarked = int(input("Enter port (0 = S, 1 = C, 2 = Q): "))
sample = [[pclass, sex, age, sibsp, parch, fare, embarked]]
prediction = tree_titanic.predict(sample)[0]
print("Prediction (1 = survived, 0 = did not survive):", prediction)
Best Practices and Troubleshooting Tips#
- Always check if there are missing values after you load data.
- Visualize features to look for outliers or surprises.
- Set random_state in train_test_split for stable results.
- Try different max_depth, min_samples_split, and split criteria.
- Avoid leaking the answer in your features.
If your tree overfits, prune it or require more samples at each branch.
Extra Decision Tree Challenges!#
- Challenge 1: Tune your tree's depth and report a test score.
- Challenge 2: Use feature_importances_ to pick the top 3 variables and build a mini-tree.
- Challenge 3: Try predicting a friend's survival on Titanic using your tree.
Bonus: Show your plot as a tree diagram, and try explaining a branch in plain language.
Recap: What Did We Learn?#
- What decision trees are and why they're great for beginners.
- The difference between ID3, C4.5, and CART.
- How to clean, prepare, and split data in pandas.
- How to fit and evaluate trees in scikit-learn.
- How to read feature importance and explain results.
Congratulations, you have now built and interpreted your own decision trees in Python!
Thanks for Joining Next Steps!#
Keep practicing building decision trees. Try out different datasets or challenges from our playlist.
If you liked this lesson, give it a like and subscribe to our YouTube channel for more hands-on data mining guides!
See you next time!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



