Lesson 13 · Statistics for data analysts
Categorical Data Analysis & Chi-Square Tests | Statistics #13
Video thirteen of the 15-part series: a deeper chi-square treatment, odds ratios, and a first bridge into logistic regression. Real mushroom specimen data:…
- CourseStatistics for data analysts
- Lesson13 of 15
- Video15 min
- FormatJupyter notebook · 12 code cells
- Data1 dataset
What you'll learn
Datasets used in this lesson
Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.
- mushroom_predictions.csv78.3 KB
📓 Full notebook
Download .ipynbStatistics for Data Analysts, Video 13: Categorical Data Analysis#
- Video thirteen of the 15-part series: a deeper chi-square treatment, odds ratios, and a first bridge into logistic regression.
- Real mushroom specimen data: physical traits, encoded as real category codes, against a real edible-or-poisonous label.
- Let's get into it.
Before You Start#
- Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
- Install statsmodels if you don't have it yet:
pip install statsmodels. - Place mushroom_predictions.csv in the same folder as this notebook.
import pandas as pd
import numpy as np
from scipy import stats
import statsmodels.formula.api as smf
mushrooms = pd.read_csv('mushroom_predictions.csv')
print(mushrooms[['odor', 'actual']].head())
print(mushrooms['actual'].value_counts())
Part 1: A Deeper Chi-Square Test of Independence#
contingency = pd.crosstab(mushrooms['odor'], mushrooms['actual'])
print(contingency)
chi2_stat, p_value, dof, expected = stats.chi2_contingency(contingency)
print(f'Real chi-square statistic: {chi2_stat:.2f}')
print(f'Real degrees of freedom: {dof}')
print(f'Real p-value: {p_value:.2e}')
expected_df = pd.DataFrame(expected, index=contingency.index, columns=contingency.columns)
print(expected_df.round(2))
Part 2: Effect Size - Cramer's V#
n = contingency.values.sum()
k = min(contingency.shape) - 1
cramers_v = np.sqrt(chi2_stat / (n * k))
print(f'Real sample size: {n}')
print(f'Real Cramer\'s V: {cramers_v:.4f}')
Part 3: Odds Ratios#
mushrooms['odor_5'] = (mushrooms['odor'] == 5).astype(int)
mushrooms['edible'] = (mushrooms['actual'] == 'e').astype(int)
two_by_two = pd.crosstab(mushrooms['odor_5'], mushrooms['edible'])
print(two_by_two)
a = two_by_two.loc[1, 1]
b = two_by_two.loc[1, 0]
c = two_by_two.loc[0, 1]
d = two_by_two.loc[0, 0]
odds_ratio = (a * d) / (b * c)
print(f'Real odds ratio: {odds_ratio:.2f}')
Part 4: The Bridge to Logistic Regression#
logit_model = smf.logit('edible ~ odor_5', data=mushrooms).fit()
print(logit_model.summary())
logit_odds_ratio = np.exp(logit_model.params['odor_5'])
print(f'Real odds ratio from part three (manual): {odds_ratio:.2f}')
print(f'Real odds ratio from logistic regression (exponentiated coefficient): {logit_odds_ratio:.2f}')
predictions = (logit_model.predict(mushrooms) > 0.5).astype(int)
accuracy = (predictions == mushrooms['edible']).mean()
print(f'Real classification accuracy using odor_5 alone: {accuracy:.1%}')
Part 5: When Chi-Square Breaks Down#
low_expected = (expected < 5).sum()
print(f'Real cells with expected count under 5: {low_expected}')
print(expected_df[expected_df < 5].dropna(how="all").dropna(axis=1, how="all"))
mushrooms['odor_4'] = (mushrooms['odor'] == 4).astype(int)
sparse_table = pd.crosstab(mushrooms['odor_4'], mushrooms['edible'])
print(sparse_table)
fisher_odds, fisher_p = stats.fisher_exact(sparse_table)
print(f'Real Fisher exact test p-value: {fisher_p:.4f}')
Wrap-Up: What You Learned#
- A full chi-square test of independence, comparing real observed and expected counts directly, on real mushroom odor and edibility data.
- Cramer's V, for measuring how strong a categorical association actually is, not just whether it's statistically real.
- Odds ratios, computed by hand from a real 2-by-2 table.
- The exact mathematical bridge from an odds ratio to logistic regression, confirmed by matching numbers, and using it to build a real, though imperfect, classifier.
- Fisher's exact test, the correct tool when real expected cell counts are too small for chi-square's approximation to hold.
- Video fourteen shifts to Bayesian statistics: priors, posteriors, and revisiting the same synthetic A/B test data from video eight through a completely different statistical lens. Subscribe so it lands automatically see you there.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



