Mathew K Analytics

Lesson 13 · Statistics for data analysts

Categorical Data Analysis & Chi-Square Tests | Statistics #13

Video thirteen of the 15-part series: a deeper chi-square treatment, odds ratios, and a first bridge into logistic regression. Real mushroom specimen data:…

What you'll learn

Datasets used in this lesson

Save these next to the notebook. In Google Colab, upload them with the 📁 icon on the left first.

📓 Full notebook

Download .ipynb

Statistics for Data Analysts, Video 13: Categorical Data Analysis#

  • Video thirteen of the 15-part series: a deeper chi-square treatment, odds ratios, and a first bridge into logistic regression.
  • Real mushroom specimen data: physical traits, encoded as real category codes, against a real edible-or-poisonous label.
  • Let's get into it.

Before You Start#

  • Open a new Jupyter Notebook in VS Code and select your Python interpreter as the kernel.
  • Install statsmodels if you don't have it yet: pip install statsmodels.
  • Place mushroom_predictions.csv in the same folder as this notebook.
import pandas as pd
import numpy as np
from scipy import stats
import statsmodels.formula.api as smf

mushrooms = pd.read_csv('mushroom_predictions.csv')
print(mushrooms[['odor', 'actual']].head())
print(mushrooms['actual'].value_counts())
   odor actual
0     2      p
1     1      p
2     5      e
3     2      p
4     2      p
actual
e    842
p    783
Name: count, dtype: int64

Part 1: A Deeper Chi-Square Test of Independence#

contingency = pd.crosstab(mushrooms['odor'], mushrooms['actual'])
print(contingency)
actual    e    p
odor            
0        82    0
1         0   38
2         0  406
3        81    0
4         0    9
5       679   23
6         0   62
7         0  126
8         0  119
chi2_stat, p_value, dof, expected = stats.chi2_contingency(contingency)
print(f'Real chi-square statistic: {chi2_stat:.2f}')
print(f'Real degrees of freedom: {dof}')
print(f'Real p-value: {p_value:.2e}')
Real chi-square statistic: 1535.90
Real degrees of freedom: 8
Real p-value: 0.00e+00
expected_df = pd.DataFrame(expected, index=contingency.index, columns=contingency.columns)
print(expected_df.round(2))
actual       e       p
odor                  
0        42.49   39.51
1        19.69   18.31
2       210.37  195.63
3        41.97   39.03
4         4.66    4.34
5       363.74  338.26
6        32.13   29.87
7        65.29   60.71
8        61.66   57.34

Part 2: Effect Size - Cramer's V#

n = contingency.values.sum()
k = min(contingency.shape) - 1
cramers_v = np.sqrt(chi2_stat / (n * k))
print(f'Real sample size: {n}')
print(f'Real Cramer\'s V: {cramers_v:.4f}')
Real sample size: 1625
Real Cramer's V: 0.9722

Part 3: Odds Ratios#

mushrooms['odor_5'] = (mushrooms['odor'] == 5).astype(int)
mushrooms['edible'] = (mushrooms['actual'] == 'e').astype(int)
two_by_two = pd.crosstab(mushrooms['odor_5'], mushrooms['edible'])
print(two_by_two)
edible    0    1
odor_5          
0       760  163
1        23  679
a = two_by_two.loc[1, 1]
b = two_by_two.loc[1, 0]
c = two_by_two.loc[0, 1]
d = two_by_two.loc[0, 0]
odds_ratio = (a * d) / (b * c)
print(f'Real odds ratio: {odds_ratio:.2f}')
Real odds ratio: 137.65

Part 4: The Bridge to Logistic Regression#

logit_model = smf.logit('edible ~ odor_5', data=mushrooms).fit()
print(logit_model.summary())
Optimization terminated successfully.
         Current function value: 0.327103
         Iterations 7
                           Logit Regression Results                           
==============================================================================
Dep. Variable:                 edible   No. Observations:                 1625
Model:                          Logit   Df Residuals:                     1623
Method:                           MLE   Df Model:                            1
Date:                Sun, 16 Aug 2026   Pseudo R-squ.:                  0.5276
Time:                        19:57:52   Log-Likelihood:                -531.54
converged:                       True   LL-Null:                       -1125.3
Covariance Type:            nonrobust   LLR p-value:                3.172e-260
==============================================================================
                 coef    std err          z      P>|z|      [0.025      0.975]
------------------------------------------------------------------------------
Intercept     -1.5396      0.086    -17.836      0.000      -1.709      -1.370
odor_5         4.9247      0.229     21.513      0.000       4.476       5.373
==============================================================================
logit_odds_ratio = np.exp(logit_model.params['odor_5'])
print(f'Real odds ratio from part three (manual): {odds_ratio:.2f}')
print(f'Real odds ratio from logistic regression (exponentiated coefficient): {logit_odds_ratio:.2f}')
Real odds ratio from part three (manual): 137.65
Real odds ratio from logistic regression (exponentiated coefficient): 137.65
predictions = (logit_model.predict(mushrooms) > 0.5).astype(int)
accuracy = (predictions == mushrooms['edible']).mean()
print(f'Real classification accuracy using odor_5 alone: {accuracy:.1%}')
Real classification accuracy using odor_5 alone: 88.6%

Part 5: When Chi-Square Breaks Down#

low_expected = (expected < 5).sum()
print(f'Real cells with expected count under 5: {low_expected}')
print(expected_df[expected_df < 5].dropna(how="all").dropna(axis=1, how="all"))
Real cells with expected count under 5: 2
actual         e         p
odor                      
4       4.663385  4.336615
mushrooms['odor_4'] = (mushrooms['odor'] == 4).astype(int)
sparse_table = pd.crosstab(mushrooms['odor_4'], mushrooms['edible'])
print(sparse_table)
fisher_odds, fisher_p = stats.fisher_exact(sparse_table)
print(f'Real Fisher exact test p-value: {fisher_p:.4f}')
edible    0    1
odor_4          
0       774  842
1         9    0
Real Fisher exact test p-value: 0.0014

Wrap-Up: What You Learned#

  • A full chi-square test of independence, comparing real observed and expected counts directly, on real mushroom odor and edibility data.
  • Cramer's V, for measuring how strong a categorical association actually is, not just whether it's statistically real.
  • Odds ratios, computed by hand from a real 2-by-2 table.
  • The exact mathematical bridge from an odds ratio to logistic regression, confirmed by matching numbers, and using it to build a real, though imperfect, classifier.
  • Fisher's exact test, the correct tool when real expected cell counts are too small for chi-square's approximation to hold.
  • Video fourteen shifts to Bayesian statistics: priors, posteriors, and revisiting the same synthetic A/B test data from video eight through a completely different statistical lens. Subscribe so it lands automatically see you there.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.