Lesson 8 · ML algorithms deep dive
Naive Bayes Classifier Explained | ML Algorithms #8
Video eight of the 12-part series: revisiting real mushroom data for another direct algorithm rematch. Real physical mushroom traits, same real dataset,…
- CourseML algorithms deep dive
- Lesson8 of 12
- Video16 min
- FormatJupyter notebook · 11 code cells
- Data1 dataset
What you'll learn
Datasets used in this lesson
Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.
- mushroom_dataset.csv78.3 KB
📓 Full notebook
Download .ipynbML Algorithms Deep-Dive, Video 8: Naive Bayes#
- Video eight of the 12-part series: revisiting real mushroom data for another direct algorithm rematch.
- Real physical mushroom traits, same real dataset, same real question, against video four's decision tree.
- Let's get into it.
Before You Start#
- Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
- You'll need pandas, NumPy, and scikit-learn.
- Place mushroom_dataset.csv in the same folder as this notebook.
import pandas as pd
import numpy as np
from sklearn.naive_bayes import CategoricalNB, GaussianNB
from sklearn.tree import DecisionTreeClassifier
from sklearn.model_selection import train_test_split, cross_val_score, StratifiedKFold
from sklearn.metrics import confusion_matrix
mushrooms = pd.read_csv('mushroom_dataset.csv')
feature_cols = [c for c in mushrooms.columns if c not in ('actual', 'predicted')]
print(f'Real mushrooms in this dataset: {len(mushrooms)}')
Part 1: Bayes' Theorem, the Core Idea#
p_poisonous = (mushrooms['actual'] == 'p').mean()
print(f'Real prior probability of poisonous: {p_poisonous:.4f}')
Part 2: Bayes' Theorem, Worked by Hand#
odor_code = 5
p_odor_given_poisonous = ((mushrooms['odor'] == odor_code) & (mushrooms['actual'] == 'p')).sum() / (mushrooms['actual'] == 'p').sum()
p_odor = (mushrooms['odor'] == odor_code).mean()
posterior = (p_odor_given_poisonous * p_poisonous) / p_odor
print(f'Real P(odor=5 | poisonous): {p_odor_given_poisonous:.4f}')
print(f'Real P(odor=5): {p_odor:.4f}')
print(f'Real P(poisonous | odor=5): {posterior:.4f}')
Part 3: One Real Feature That Almost Solves It#
for code in sorted(mushrooms['odor'].unique()):
subset = mushrooms[mushrooms['odor'] == code]
real_rate = (subset['actual'] == 'p').mean()
print(f'odor={code}: real n={len(subset)}, real P(poisonous|odor)={real_rate:.4f}')
Part 4: Fitting a Real Naive Bayes Model#
X = mushrooms[feature_cols].values
y = mushrooms['actual'].values
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, random_state=42, stratify=y)
categorical_nb = CategoricalNB().fit(X_train, y_train)
print(f'Real CategoricalNB test accuracy: {categorical_nb.score(X_test, y_test):.4f}')
Part 5: Choosing the Right Naive Bayes Variant#
gaussian_nb = GaussianNB().fit(X_train, y_train)
print(f'Real GaussianNB test accuracy (features misassumed continuous): {gaussian_nb.score(X_test, y_test):.4f}')
Part 6: The Conditional Independence Assumption#
poisonous_only = mushrooms[mushrooms['actual'] == 'p']
edible_only = mushrooms[mushrooms['actual'] == 'e']
corr_poisonous = poisonous_only['gill-color'].corr(poisonous_only['spore-print-color'])
corr_edible = edible_only['gill-color'].corr(edible_only['spore-print-color'])
print(f'Real correlation between gill color and spore print color, within poisonous mushrooms: {corr_poisonous:.4f}')
print(f'Real correlation between gill color and spore print color, within edible mushrooms: {corr_edible:.4f}')
That's a real, direct violation of Naive Bayes' core assumption, at least within the poisonous group: gill color and spore print color aren't independent once you already know a mushroom is poisonous, they move together. Naive Bayes still runs the real math as if they were independent anyway, and that's exactly where its real accuracy has room to suffer.
Part 7: The Real Rematch#
skf = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
nb_cv = cross_val_score(CategoricalNB(), X_train, y_train, cv=skf, scoring='accuracy')
tree_cv = cross_val_score(DecisionTreeClassifier(max_depth=5, random_state=42), X_train, y_train, cv=skf, scoring='accuracy')
print(f'Real Naive Bayes mean CV accuracy: {nb_cv.mean():.4f}')
print(f'Real decision tree (depth 5) mean CV accuracy: {tree_cv.mean():.4f}')
Part 8: Why the Gap Actually Matters Here#
nb_final = CategoricalNB().fit(X_train, y_train)
nb_pred = nb_final.predict(X_test)
cm = confusion_matrix(y_test, nb_pred, labels=['e', 'p'])
print('Real Naive Bayes confusion matrix (rows=actual edible/poisonous, cols=predicted edible/poisonous):')
print(cm)
print(f'Real poisonous mushrooms misclassified as edible: {cm[1, 0]}')
Part 9: Strengths, Weaknesses, and When to Use It#
- Strength: genuinely fast to train, even on real datasets with many features.
- Strength: works well with surprisingly little real training data, since it just counts real co-occurrences.
- Weakness: the real independence assumption actively hurts accuracy whenever real features genuinely interact, as parts six through eight showed directly.
- Weakness: choosing the wrong real variant for your feature types, as part five demonstrated, is an easy real mistake to make.
- Use it as a genuinely fast real baseline, especially for high-dimensional real data like text, but verify the independence assumption before trusting it in a genuinely high-stakes setting.
Part 10: Saving Your Work#
import matplotlib.pyplot as plt
plt.figure(figsize=(6, 5))
plt.imshow(cm, cmap='Reds')
plt.xticks([0, 1], ['Predicted edible', 'Predicted poisonous'])
plt.yticks([0, 1], ['Actual edible', 'Actual poisonous'])
for i in range(2):
for j in range(2):
plt.text(j, i, str(cm[i, j]), ha='center', va='center', fontsize=14)
plt.title('Real Naive Bayes Confusion Matrix')
plt.colorbar()
plt.savefig('naive_bayes_confusion_matrix.png', dpi=150)
plt.show()
import joblib
joblib.dump(nb_final, 'naive_bayes_model.joblib')
reloaded_nb = joblib.load('naive_bayes_model.joblib')
print(f'Real original model test accuracy: {nb_final.score(X_test, y_test):.4f}')
print(f'Real reloaded model test accuracy: {reloaded_nb.score(X_test, y_test):.4f}')
Wrap-Up: What You Learned#
- Bayes' theorem, worked by hand on a real mushroom trait, turning a real prior into a real posterior.
- One real feature, odor, that came remarkably close to solving the whole problem on its own.
- Fitting real Naive Bayes, and picking the correct real variant for categorical data.
- Testing the conditional independence assumption directly, and finding a real, substantial violation.
- A real rematch against video four's decision tree, with the accuracy gap traced back to that real violation.
- Why that gap translates into a genuinely costly kind of real mistake in a safety-relevant task.
- Video nine turns to unsupervised learning: K-Means clustering, on real, unlabeled mall customer data. Subscribe so it lands automatically see you there.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



