Mathew K Analytics

Lesson 17 · Data Mining

Mastering Naive Bayes Classifier: Key Concepts and Real-World Data Analysis Applications

Welcome to our hands-on lesson focusing on the Naive Bayes algorithm for classification problems. In this lesson, you will: Learn what Naive Bayes is and…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Week 56: Introduction to Naive Bayes Classifier#

Welcome to our hands-on lesson focusing on the Naive Bayes algorithm for classification problems.

In this lesson, you will:

  • Learn what Naive Bayes is and why it is useful.
  • Explore data cleaning and preparation.
  • Apply Naive Bayes to real-world telecom churn data.
  • Practice predictive modeling and evaluation.
  • Try classification exercises and discover best-practices.
# Data setup (Telecom Customer Churn Dataset)
import warnings; warnings.filterwarnings("ignore")
import pandas as pd
import numpy as np
np.random.seed(42)
url = 'https://raw.githubusercontent.com/IBM/telco-customer-churn-on-icp4d/master/data/Telco-Customer-Churn.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
(7043, 21)
   customerID  gender  SeniorCitizen Partner Dependents  tenure PhoneService  \
0  7590-VHVEG  Female              0     Yes         No       1           No   
1  5575-GNVDE    Male              0      No         No      34          Yes   
2  3668-QPYBK    Male              0      No         No       2          Yes   

      MultipleLines InternetService OnlineSecurity  ... DeviceProtection  \
0  No phone service             DSL             No  ...               No   
1                No             DSL            Yes  ...              Yes   
2                No             DSL            Yes  ...               No   

  TechSupport StreamingTV StreamingMovies        Contract PaperlessBilling  \
0          No          No              No  Month-to-month              Yes   
1          No          No              No        One year               No   
2          No          No              No  Month-to-month              Yes   

      PaymentMethod MonthlyCharges  TotalCharges Churn  
0  Electronic check          29.85         29.85    No  
1      Mailed check          56.95        1889.5    No  
2      Mailed check          53.85        108.15   Yes  

[3 rows x 21 columns]

What is Naive Bayes?#

Naive Bayes is a fast classification algorithm based on Bayes Theorem.

It calculates the chance of each outcome using simple probabilities.

It is called 'naive' because it assumes all input features are independent of each other.

Naive Bayes is a favorite for text classification and easy risk prediction.

# Let us look for missing values
missing_counts = df.isnull().sum()
print(missing_counts[missing_counts > 0])
Series([], dtype: int64)
# Inspect column types
print(df.dtypes)
customerID           object
gender               object
SeniorCitizen         int64
Partner              object
Dependents           object
tenure                int64
PhoneService         object
MultipleLines        object
InternetService      object
OnlineSecurity       object
OnlineBackup         object
DeviceProtection     object
TechSupport          object
StreamingTV          object
StreamingMovies      object
Contract             object
PaperlessBilling     object
PaymentMethod        object
MonthlyCharges      float64
TotalCharges         object
Churn                object
dtype: object
# Convert TotalCharges to numeric and handle errors
df['TotalCharges'] = pd.to_numeric(df['TotalCharges'], errors='coerce')
print(df['TotalCharges'].isnull().sum())
11
# Remove rows with missing values for simplicity
df = df.dropna()
print(df.shape)
(7032, 21)

Preparing for Naive Bayes#

Naive Bayes can only work with numbers, not words.

We will convert all categorical columns (like yes/no) into new number columns.

We will also keep 'Churn' as our label for prediction.

# Turn yes/no columns into numbers
for col in ['Partner', 'Dependents', 'PhoneService', 'PaperlessBilling', 'Churn']:
    df[col] = df[col].map({'Yes':1, 'No':0})
print(df[['Partner','Dependents','PhoneService','PaperlessBilling','Churn']].head(3))
   Partner  Dependents  PhoneService  PaperlessBilling  Churn
0        1           0             0                 1      0
1        0           0             1                 0      0
2        0           0             1                 1      1
# Turn remaining strings into numbers with one-hot encoding
df_encoded = pd.get_dummies(df, drop_first=True)
print(df_encoded.shape)
print(df_encoded.head(3))
(7032, 7062)
   SeniorCitizen  Partner  Dependents  tenure  PhoneService  PaperlessBilling  \
0              0        1           0       1             0                 1   
1              0        0           0      34             1                 0   
2              0        0           0       2             1                 1   

   MonthlyCharges  TotalCharges  Churn  customerID_0003-MKNFE  ...  \
0           29.85         29.85      0                  False  ...   
1           56.95       1889.50      0                  False  ...   
2           53.85        108.15      1                  False  ...   

   TechSupport_Yes  StreamingTV_No internet service  StreamingTV_Yes  \
0            False                            False            False   
1            False                            False            False   
2            False                            False            False   

   StreamingMovies_No internet service  StreamingMovies_Yes  \
0                                False                False   
1                                False                False   
2                                False                False   

   Contract_One year  Contract_Two year  \
0              False              False   
1               True              False   
2              False              False   

   PaymentMethod_Credit card (automatic)  PaymentMethod_Electronic check  \
0                                  False                            True   
1                                  False                           False   
2                                  False                           False   

   PaymentMethod_Mailed check  
0                       False  
1                        True  
2                        True  

[3 rows x 7062 columns]

Exploratory Data Analysis (EDA)#

It helps to look at the balance of classes in our label column.

We will plot the Churn column to see how many customers left versus stayed.

# Plot distribution of Churn
import matplotlib.pyplot as plt
df['Churn'].value_counts().plot(kind='bar')
plt.title('Customer Churn Distribution')
plt.xlabel('Churn (0=No, 1=Yes)')
plt.ylabel('Number of Customers')
plt.show()
No description has been provided for this image
# Select features and label
X = df_encoded.drop('Churn', axis=1)
y = df_encoded['Churn']
print(X.shape, y.shape)
(7032, 7061) (7032,)
# Split data into train and test sets
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)
print(X_train.shape, X_test.shape)
print(y_train.shape, y_test.shape)
(4922, 7061) (2110, 7061)
(4922,) (2110,)
# Train a Naive Bayes classifier
from sklearn.naive_bayes import GaussianNB
nb = GaussianNB()
nb.fit(X_train, y_train)
GaussianNB()
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
# Predict on test data
y_pred = nb.predict(X_test)
print(y_pred[:10])
[0 0 1 1 1 1 1 1 1 0]
# Check prediction accuracy
from sklearn.metrics import accuracy_score
accuracy = accuracy_score(y_test, y_pred)
print(f"Accuracy: {accuracy:.2f}")
Accuracy: 0.60
# Print a detailed classification report
from sklearn.metrics import classification_report
print(classification_report(y_test, y_pred))
              precision    recall  f1-score   support

           0       0.94      0.48      0.64      1549
           1       0.39      0.92      0.55       561

    accuracy                           0.60      2110
   macro avg       0.67      0.70      0.59      2110
weighted avg       0.80      0.60      0.62      2110

Try it yourself: Predict a random customer#

Let us try out the model manually. Enter example values for a random customer's features.

# Manual prediction with user input
features = list(X.columns)
sample = []
for f in features[:3]:
    val = float(input(f"Enter value for {f}: "))
    sample.append(val)
sample += [0]*(len(features)-3)
import numpy as np
result = nb.predict(np.array(sample).reshape(1, -1))
print(f"Predicted churn: {result[0]}")
Predicted churn: 1
# (Optional) Compare Naive Bayes to Decision Tree
from sklearn.tree import DecisionTreeClassifier
tree = DecisionTreeClassifier(random_state=42)
tree.fit(X_train, y_train)
y_tree = tree.predict(X_test)
print(f"Tree accuracy: {accuracy_score(y_test, y_tree):.2f}")
Tree accuracy: 0.77

Best Practices and Troubleshooting#

  • Always clean and preprocess your data for missing or weird values.
  • Check class balance before training.
  • Try several types of models, not just Naive Bayes.
  • Tune model parameters using validation splits.
  • Do not trust accuracy alone with imbalanced data.
  • Always test your model with new data!

Extra Tips#

  • Naive Bayes is fast, so it is great for large or messy data.
  • It works even if your features have many categories.
  • Use with text for email spam or social media classification.
  • Use with patient data for simple health risk alerts.
# Challenge: Use Naive Bayes on Titanic data
import pandas as pd
url = 'https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv'
titanic = pd.read_csv(url)
print(titanic.shape)
print(titanic.head(2))
(891, 12)
   PassengerId  Survived  Pclass  \
0            1         0       3   
1            2         1       1   

                                                Name     Sex   Age  SibSp  \
0                            Braund, Mr. Owen Harris    male  22.0      1   
1  Cumings, Mrs. John Bradley (Florence Briggs Th...  female  38.0      1   

   Parch     Ticket     Fare Cabin Embarked  
0      0  A/5 21171   7.2500   NaN        S  
1      0   PC 17599  71.2833   C85        C  

Recap and Next Steps#

You learned how to prepare real-world data, fit a Naive Bayes classifier, and check your model's accuracy.

Practicing with new datasets makes you a stronger data miner.

Try re-running parts of this lesson. Explore mistakes and new ideas!

Thanks for Learning! Like & Subscribe for More#

If you found value in this lesson, please hit the like button and subscribe. This helps me create more beginner Python tutorials just for you.

See you next week for more data mining hands-on lessons!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.