Lesson 17 · Data Mining
Mastering Naive Bayes Classifier: Key Concepts and Real-World Data Analysis Applications
Welcome to our hands-on lesson focusing on the Naive Bayes algorithm for classification problems. In this lesson, you will: Learn what Naive Bayes is and…
- CourseData Mining
- Lesson17 of 31
- Video29 min
- FormatJupyter notebook · 17 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbWeek 56: Introduction to Naive Bayes Classifier#
Welcome to our hands-on lesson focusing on the Naive Bayes algorithm for classification problems.
In this lesson, you will:
- Learn what Naive Bayes is and why it is useful.
- Explore data cleaning and preparation.
- Apply Naive Bayes to real-world telecom churn data.
- Practice predictive modeling and evaluation.
- Try classification exercises and discover best-practices.
# Data setup (Telecom Customer Churn Dataset)
import warnings; warnings.filterwarnings("ignore")
import pandas as pd
import numpy as np
np.random.seed(42)
url = 'https://raw.githubusercontent.com/IBM/telco-customer-churn-on-icp4d/master/data/Telco-Customer-Churn.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
What is Naive Bayes?#
Naive Bayes is a fast classification algorithm based on Bayes Theorem.
It calculates the chance of each outcome using simple probabilities.
It is called 'naive' because it assumes all input features are independent of each other.
Naive Bayes is a favorite for text classification and easy risk prediction.
# Let us look for missing values
missing_counts = df.isnull().sum()
print(missing_counts[missing_counts > 0])
# Inspect column types
print(df.dtypes)
# Convert TotalCharges to numeric and handle errors
df['TotalCharges'] = pd.to_numeric(df['TotalCharges'], errors='coerce')
print(df['TotalCharges'].isnull().sum())
# Remove rows with missing values for simplicity
df = df.dropna()
print(df.shape)
Preparing for Naive Bayes#
Naive Bayes can only work with numbers, not words.
We will convert all categorical columns (like yes/no) into new number columns.
We will also keep 'Churn' as our label for prediction.
# Turn yes/no columns into numbers
for col in ['Partner', 'Dependents', 'PhoneService', 'PaperlessBilling', 'Churn']:
df[col] = df[col].map({'Yes':1, 'No':0})
print(df[['Partner','Dependents','PhoneService','PaperlessBilling','Churn']].head(3))
# Turn remaining strings into numbers with one-hot encoding
df_encoded = pd.get_dummies(df, drop_first=True)
print(df_encoded.shape)
print(df_encoded.head(3))
Exploratory Data Analysis (EDA)#
It helps to look at the balance of classes in our label column.
We will plot the Churn column to see how many customers left versus stayed.
# Plot distribution of Churn
import matplotlib.pyplot as plt
df['Churn'].value_counts().plot(kind='bar')
plt.title('Customer Churn Distribution')
plt.xlabel('Churn (0=No, 1=Yes)')
plt.ylabel('Number of Customers')
plt.show()
# Select features and label
X = df_encoded.drop('Churn', axis=1)
y = df_encoded['Churn']
print(X.shape, y.shape)
# Split data into train and test sets
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)
print(X_train.shape, X_test.shape)
print(y_train.shape, y_test.shape)
# Train a Naive Bayes classifier
from sklearn.naive_bayes import GaussianNB
nb = GaussianNB()
nb.fit(X_train, y_train)
# Predict on test data
y_pred = nb.predict(X_test)
print(y_pred[:10])
# Check prediction accuracy
from sklearn.metrics import accuracy_score
accuracy = accuracy_score(y_test, y_pred)
print(f"Accuracy: {accuracy:.2f}")
# Print a detailed classification report
from sklearn.metrics import classification_report
print(classification_report(y_test, y_pred))
Try it yourself: Predict a random customer#
Let us try out the model manually. Enter example values for a random customer's features.
# Manual prediction with user input
features = list(X.columns)
sample = []
for f in features[:3]:
val = float(input(f"Enter value for {f}: "))
sample.append(val)
sample += [0]*(len(features)-3)
import numpy as np
result = nb.predict(np.array(sample).reshape(1, -1))
print(f"Predicted churn: {result[0]}")
# (Optional) Compare Naive Bayes to Decision Tree
from sklearn.tree import DecisionTreeClassifier
tree = DecisionTreeClassifier(random_state=42)
tree.fit(X_train, y_train)
y_tree = tree.predict(X_test)
print(f"Tree accuracy: {accuracy_score(y_test, y_tree):.2f}")
Best Practices and Troubleshooting#
- Always clean and preprocess your data for missing or weird values.
- Check class balance before training.
- Try several types of models, not just Naive Bayes.
- Tune model parameters using validation splits.
- Do not trust accuracy alone with imbalanced data.
- Always test your model with new data!
Extra Tips#
- Naive Bayes is fast, so it is great for large or messy data.
- It works even if your features have many categories.
- Use with text for email spam or social media classification.
- Use with patient data for simple health risk alerts.
# Challenge: Use Naive Bayes on Titanic data
import pandas as pd
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
titanic = pd.read_csv(url)
print(titanic.shape)
print(titanic.head(2))
Recap and Next Steps#
You learned how to prepare real-world data, fit a Naive Bayes classifier, and check your model's accuracy.
Practicing with new datasets makes you a stronger data miner.
Try re-running parts of this lesson. Explore mistakes and new ideas!
Thanks for Learning! Like & Subscribe for More#
If you found value in this lesson, please hit the like button and subscribe. This helps me create more beginner Python tutorials just for you.
See you next week for more data mining hands-on lessons!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



